List of exercises

Full list

This is a list of all exercises and solutions in this lesson, mainly as a reference for helpers and instructors. This list is automatically generated from all of the other pages in the lesson. Any single teaching event will probably cover only a subset of these, depending on their interests.

Social coding

In social-coding.md:

Social-1: Think about if and how you share

  • Did you ever share your code? If yes, what motivated you? Come up with reasons for sharing your scripts/code/data.

  • Also think about reasons for not sharing.

In social-coding.md:

Social-2: Discussion about “You aren’t required to support anyone”

  • Have you experienced an implicit expectation of support?

  • Supporting all requests can lead to overworking and mental health issues.

  • Not supporting requests can also induce guilt.

  • Most projects are maintained by 1 or 2 persons.

  • Most projects cannot retain contributors for a longer time. Interests change. “Casual contributors are like tourists visiting NYC for a weekend” (Nadia Asparouhova, book below).

  • If you maintain all projects that you start forever, at some point it may be difficult to start new projects.

  • What are your experiences? Do you agree with the above thoughts?

  • Book recommendation: Nadia Asparouhova (formerly Nadia Eghbal): “Working in Public: The Making and Maintenance of Open Source Software (Stripe Press)”

Software licensing focusing on open source

In software-licensing.md:

Limitations and context of this lesson

This lesson is designed as practical educational material for researchers and research software engineers, not formal legal advice.

  • EU directives set only minimum requirements in some areas: Member States implement them differently and may add national rules not covered here. For example, some Member States let university researchers retain ownership of the programs they write instead of applying the employer rule in Art. 2(3).

  • Institutional Context: Employment contracts, grant agreements, and university policies heavily influence software ownership and licensing choices.

  • This lesson covers only the general principles of open-source reuse, copyright scope, and software adaptation.

If you need formal guidance, the references below can help and so can legal experts, especially if your host institute has a legal services office:

In software-licensing.md:

Licensing-1: Which scenarios typically describe your work?

The text below can be copied to the collaborative document for an online poll:

## Question: How do you work with other's code?

**Choose many**. Vote by adding an `o` character:

- 1. Writing everything yourself
  - votes:

- 2. Implementing a published algorithm 
  - votes:

- 3. Pasting in a permissive snippet
  - votes:

- 4. Pasting in a copyleft snippet
  - votes:

- 5. Importing or linking a library
  - votes:

- 6. Writing a Dockerfile or .def
  - votes:

- 7. Publishing a built image
  - votes:

- 8. Using Copilot, ChatGPT, Claude, or similar
  - votes:

- 9. Shipping prompts, weights or datasets
  - votes:

In software-licensing.md:

Licensing-2: Can you decide from the number of lines?

The text below can be copied to the collaborative document for an online poll:

## Question: Which of these can you safely copy *based only on its size*?

**Choose many**. Vote by adding an `o` character:

- A. A one-line expression: `return max(lo, min(x, hi))`
  - votes:

- B. Five lines of ordinary boilerplate for parsing command-line arguments
  - votes:

- C. Three unusually written lines copied verbatim from a GPL-licensed solver
  - votes:

- D. Twenty lines you wrote independently after reading an algorithm in a paper, without looking at another implementation
  - votes:

- E. Anything under 10 lines is too small to be copyrighted
  - votes:

- F. None of the above: the number of lines alone does not decide
  - votes:


### Follow-up question

For each example, what information would you want to know before reusing or publishing the code?

In software-licensing.md:

Scenario 1: Authoring original code and algorithms

You wrote an original algorithm from scratch (in Python, C++, Rust, etc.). Your repository contains only your original source code and dependency specifications (requirements.txt, CMakeLists.txt, Cargo.toml).

  • Licensing Goal: You want maximum adoption and zero friction for commercial or academic reuse.

  • Legal Reality: External dependencies remain separate works. Because you only list them and have not bundled third-party code inside your repository, no inbound license terms constrain your choice. This changes if you ship dependencies together with your code, for example in an executable or container image (see Scenario 5 and Scenario 7).

  • JLA Selection Strategy: To ensure downstream users must keep your copyright notice while granting them maximum flexibility to incorporate your code into both open and proprietary software, you require Incl. Copyright without imposing share-alike conditions (leaving Copyleft/Share a. unselected).

In software-licensing.md:

Scenario 2: Choosing reciprocity for your own implementation

You developed a custom solver implementing algorithms from academic literature. You want any downstream improvements, extensions, or modifications to remain open-source and be shared back with the scientific community.

  • Licensing Goal: You want to enforce reciprocity (share-alike), preventing third parties from incorporating your implementation into proprietary software without sharing their modifications.

  • Legal Reality: The published algorithm itself is an unprotected idea anyone may implement it independently, as Scenario 1 and the SAS ruling establish. What copyright protects is your specific implementation. Nothing about implementing a published method forces a particular license; copyleft here is your deliberate choice to bind downstream distributors to matching terms.

  • JLA Selection Strategy: To enforce reciprocal sharing, you must mandate that downstream distributors disclose their modified source code (Disclose source) and license their adaptations or combined works under matching terms (Copyleft/Share a.).

In software-licensing.md:

Scenario 3: Embedding permissively licensed third-party code

You are building an RSE tool and copied a helper function or utility snippet from a third-party project licensed under a permissive license (e.g., MIT or Apache-2.0) directly into one of your source files.

  • Licensing Goal: You want to maintain a permissive default for your project while properly acknowledging and legally respecting the embedded third-party code.

  • Legal Reality: Permissive licenses explicitly grant you permission to copy, modify, and embed their code into your repository. However, embedding permissive code does not make the original third-party copyright disappear: you must preserve the original copyright notice and license terms for that specific snippet.

  • JLA Selection Strategy: Because inbound permissive code gives you maximum licensing flexibility, your overall repository can remain permissively licensed. To reflect this, require that notices are kept (Incl. Copyright) without imposing reciprocal sharing constraints (leaving Copyleft/Share a. unselected).

In software-licensing.md:

Scenario 4: Embedding copyleft third-party code

You are building a software tool and copied a non-trivial code snippet from a third-party project licensed under a copyleft license (e.g., GPL-3.0 or EUPL-1.2) directly into one of your source files. This is the situation behind the failed build (job #142) in the Motivation section.

  • Licensing Goal: Comply with legal requirements imposed by the inbound copyleft code while ensuring your overall repository remains legally compliant.

  • Legal Reality: Copying a non-trivial copyleft snippet into your source files creates a single combined work, so copyleft licensing generally extends to your whole project. Moving the snippet into a separate file of the same program does not change this. “Non-trivial” matters: a snippet too short or purely functional to qualify as the author’s own intellectual creation (Art. 1(3)) may not carry copyright at all. There is no word count or line count that draws this line if you are unsure, assume it is protected and either comply or reimplement.

  • JLA Selection Strategy: If you keep the snippet, your repository must adopt reciprocal sharing terms, so configure JLA to require source code disclosure (Disclose source) and reciprocal licensing (Copyleft/Share a.).

In software-licensing.md:

Scenario 5: Linking against a GPL-licensed library

You are developing a software application that imports or links against an external software library licensed under GPL-3.0 (e.g., importing a GPL Python package or linking a C/C++ static/shared library).

  • Licensing Goal: Ensure legal compliance while using copyleft libraries as core dependencies in your software project.

  • Legal Reality: Whether linking creates a combined work is genuinely unsettled, and often has to be decided case by case. The FSF’s position is that linking a GPL library statically or dynamically creates a combined work; some legal scholars and Commission EUPL guidance disagree, particularly for dynamic linking through a stable API. Most Member States have no case law on this, so no firm general rule can be stated. The guidance below follows the conservative, widely-adopted reading.

  • JLA Selection Strategy: Under the conservative reading, linking to a GPL library means the combined program you distribute must be released under matching reciprocal terms, so configure JLA to require source code disclosure (Disclose source) and reciprocal licensing (Copyleft/Share a.).

In software-licensing.md:

Scenario 6: Authoring container recipes and environment specifications

You are creating a Dockerfile, Apptainer .def file, Conda environment.yml, or build recipe to automate the setup of your research environment. The recipe itself contains setup instructions, shell commands, and package lists.

  • Licensing Goal: You want maximum adoption and reuse of your build automation script so other researchers can freely adapt and build upon your workflow.

  • Legal Reality: Build recipes and configuration scripts are plain-text source code, separate from the software binaries they download at build time. The build instructions you write are your expression but note that a very short recipe (a FROM line plus two RUN commands) may be too trivial to meet the Art. 1(3) originality threshold and may not attract copyright at all. The same applies to a plain list of package names in an environment.yml. Longer, non-obvious recipes clearly do attract copyright.

  • JLA Selection Strategy: To allow anyone to reuse or adapt your container recipe without restrictions, require that your copyright notice is kept (Incl. Copyright) while leaving reciprocal requirements (Copyleft/Share a.) unselected.

In software-licensing.md:

Scenario 7: Distributing pre-built container images

You compiled and published a pre-built container image (e.g., pushing a compiled Docker image to Docker Hub, GitHub Container Registry, or an institutional registry, or sharing an Apptainer .sif file) containing an OS layer, runtime binaries, dependencies, and your application code.

  • Licensing Goal: Safely distribute compiled container images without violating the license terms of any software layer or binary included inside the image.

  • Legal Reality: A compiled container image is a multi-license aggregate bundle, not a single combined work. Distributing pre-built binaries makes you a distributor of every package inside, so source-availability obligations apply to the copyleft components (Linux base packages, coreutils, GPL libraries). But those packages sitting in the same filesystem as your application do not make your application a derivative of them: this is mere aggregation. Your own code keeps whatever license you chose; you simply also carry distributor obligations for the copyleft software you are shipping alongside it.

  • JLA Selection Strategy: Because a container image combines multiple distinct software components, JLA is used to evaluate constituent component obligations. When distributing compiled binaries containing copyleft layers, source disclosure requirements (Disclose source) must be fulfilled for those specific layers.

In software-licensing.md:

Scenario 8: AI-assisted code generation

You used AI tools (e.g., GitHub Copilot, ChatGPT, Claude) to write functions, unit tests, or documentation for your research software repository.

  • Licensing Goal: Apply a permissive license (MIT or Apache-2.0) to your repository with confidence, without incurring hidden copyright infringement or copyleft obligations from code the AI model reproduced from its training data.

  • Legal Reality: Unmodified AI-generated outputs lack human authorship and are generally not eligible for copyright protection under current EU and international legal standards. Most real code, however, is a mix of human and AI contribution: you prompt, select, edit, and integrate. Where the line falls between AI output and your own work is unsettled and varies between Member States; there is no percentage or line-count threshold. The more you design, choose, edit, and integrate, the stronger your claim that the result is your work. Separately, if an LLM reproduces a substantial copyrighted code snippet verbatim from its training data (memorization), that output snippet retains its original copyright and license obligations.

  • JLA Selection Strategy: To ensure maximum adoption and academic reuse for your overall codebase, require that your copyright notice is kept (Incl. Copyright) while avoiding share-alike constraints (leaving Copyleft/Share a. unselected), supported by automated compliance checks.

In software-licensing.md:

Scenario 9: Packaging AI workflows, datasets, and model weights

You are developing research software that includes source code alongside trained machine learning model weights (.pt, .safetensors) and benchmark datasets.

  • Licensing Goal: Apply a clear licensing structure, with different licenses for different parts of the repository, that makes both the software source code and the non-code assets (data, weights) open and reusable under appropriate legal frameworks.

  • Legal Reality: Standard software licenses (MIT, GPL) are written for source code and fit datasets and model parameters poorly. Datasets may attract the EU sui generis database right where there has been substantial investment in obtaining, verifying, or presenting their contents. Model weights are a harder case: they are the numerical values learned during training, neither code nor a database, and whether they attract any copyright protection in the EU is genuinely unsettled. Because of this uncertainty, applying an explicit license to weights is about setting clear terms for your users, not about relying on a settled legal right.

  • JLA Selection Strategy: Use JLA to select an OSI-approved open-source license for the executable code component (Incl. Copyright selected), while using Creative Commons licenses (e.g., CC-BY-4.0 or CC0-1.0) for the dataset and weight files.

Software citation

In software-citation.md:

Discussion (Citation-1): Explain how you currently cite software

  • Do you cite software that you use? How?

  • If I wanted to cite your code/scripts, what would I need to do?