List of exercises
Full list
This is a list of all exercises and solutions in this lesson, mainly as a reference for helpers and instructors. This list is automatically generated from all of the other pages in the lesson. Any single teaching event will probably cover only a subset of these, depending on their interests.
Software licensing focusing on open source
Limitations and context of this lesson
This lesson is designed as practical educational material for researchers and research software engineers, not formal legal advice.
EU directives set only minimum requirements in some areas: Member States implement them differently and may add national rules not covered here. For example, some Member States let university researchers retain ownership of the programs they write instead of applying the employer rule in Art. 2(3).
Institutional Context: Employment contracts, grant agreements, and university policies heavily influence software ownership and licensing choices.
This lesson covers only the general principles of open-source reuse, copyright scope, and software adaptation.
If you need formal guidance, the references below can help and so can legal experts, especially if your host institute has a legal services office:
Licensing-1: Which scenarios typically describe your work?
The text below can be copied to the collaborative document for an online poll:
## Question: How do you work with other's code?
**Choose many**. Vote by adding an `o` character:
- 1. Writing everything yourself
- votes:
- 2. Implementing a published algorithm
- votes:
- 3. Pasting in a permissive snippet
- votes:
- 4. Pasting in a copyleft snippet
- votes:
- 5. Importing or linking a library
- votes:
- 6. Writing a Dockerfile or .def
- votes:
- 7. Publishing a built image
- votes:
- 8. Using Copilot, ChatGPT, Claude, or similar
- votes:
- 9. Shipping prompts, weights or datasets
- votes:
Licensing-2: Can you decide from the number of lines?
The text below can be copied to the collaborative document for an online poll:
## Question: Which of these can you safely copy *based only on its size*?
**Choose many**. Vote by adding an `o` character:
- A. A one-line expression: `return max(lo, min(x, hi))`
- votes:
- B. Five lines of ordinary boilerplate for parsing command-line arguments
- votes:
- C. Three unusually written lines copied verbatim from a GPL-licensed solver
- votes:
- D. Twenty lines you wrote independently after reading an algorithm in a paper, without looking at another implementation
- votes:
- E. Anything under 10 lines is too small to be copyrighted
- votes:
- F. None of the above: the number of lines alone does not decide
- votes:
### Follow-up question
For each example, what information would you want to know before reusing or publishing the code?
Solution
The key answer is F: there is no fixed safe number of lines.
Copyright does not use a numerical threshold such as 5, 10, or 20 lines. The important question is whether what has been copied is protected expression. Very short or purely functional code may not meet the originality threshold, while a short but distinctive piece of code may.
A and B: they may be too simple, conventional, or constrained by function to contain protectable expression, but their size alone does not answer the question.
C: being only three lines does not automatically make copied code unprotected. Check its provenance and license.
D: independently implementing the idea or algorithm is different from copying somebody else’s expression of it.
E: there is no “10-line rule”.
Practical rule: if you copied code and are unsure whether it is protected, check where it came from and under which license it was published. Preserve any required notices, or independently implement the underlying idea instead of copying the code.
Scenario 2: Choosing reciprocity for your own implementation
You developed a custom solver implementing algorithms from academic literature. You want any downstream improvements, extensions, or modifications to remain open-source and be shared back with the scientific community.
Licensing Goal: You want to enforce reciprocity (share-alike), preventing third parties from incorporating your implementation into proprietary software without sharing their modifications.
Legal Reality: The published algorithm itself is an unprotected idea anyone may implement it independently, as Scenario 1 and the SAS ruling establish. What copyright protects is your specific implementation. Nothing about implementing a published method forces a particular license; copyleft here is your deliberate choice to bind downstream distributors to matching terms.
JLA Selection Strategy: To enforce reciprocal sharing, you must mandate that downstream distributors disclose their modified source code (
Disclose source) and license their adaptations or combined works under matching terms (Copyleft/Share a.).
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. Copyright,Disclose source, andCopyleft/Share a.Support Column: Select
OSI approved
Example JLA Matches:
EUPL-1.2,GPL-3.0,AGPL-3.0Copyleft Mechanics (EUPL vs. GPL Nuance):
GPL-3.0is the standard global copyleft license, butEUPL-1.2is specifically tailored for European institutions. EUPL-1.2 is officially published in 23 EU language versions (each with equal legal validity) and sets the applicable law and courts by reference to the licensor’s EU Member State. Its compatibility clause works in one direction only: EUPL code can be combined into a GPL project and distributed under GPL, but GPL code cannot be re-licensed under EUPL.AGPL and network use:
AGPL-3.0adds one rule to GPL-3.0: if you modify the software and let people use it over a network (for example a web portal or API), you must offer them the source code, even if you never distribute copies. Consider it if your group runs research software as an online service.The trade-offs of strong copyleft: Reciprocity comes at a cost. Some companies and projects avoid copyleft code entirely, so you may reach fewer users than with a permissive license. Strong copyleft licenses are also frequently incompatible with each other, so a future collaborator on a differently-licensed copyleft project may be unable to use your work at all (Scenario 5 covers this). Neither choice is better: Scenario 1 optimizes for reach, this scenario for keeping improvements open.
Choose deliberately, early: Changing your license later is only possible if you hold all the rights. Once others have contributed code, you need every contributor’s agreement to re-license.
Downstream Obligations: Anyone who distributes your code or a modified version of it must provide complete access to the corresponding source code under the same copyleft license and preserve your original copyright notices. Running modified code internally, without distributing it, creates no obligation (except under AGPL for network use).
Allowed Inbound Snippets: You can freely embed code snippets licensed under permissive terms (e.g., MIT, Apache-2.0, BSD), keeping their notices, or public domain waivers (CC0). Copyleft snippets must be compatible with your license, and compatibility is directional: a GPL project can take EUPL or GPL snippets, but a GPL snippet in an EUPL project would require the combined work to be distributed under GPL. When in doubt, only embed copyleft code under the same license as your project. You cannot embed closed-source or proprietary code snippets.
In-File Identification (SPDX): Apply standard machine-readable SPDX tags directly at the top of your scripts, following the REUSE specification. If you choose GPL, use
GPL-3.0-onlyorGPL-3.0-or-laterrather than plainGPL-3.0, since the two behave differently when a new GPL version is released:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: EUPL-1.2
import numpy as np
Scenario 5: Linking against a GPL-licensed library
You are developing a software application that imports or links against an external software library licensed under GPL-3.0 (e.g., importing a GPL Python package or linking a C/C++ static/shared library).
Licensing Goal: Ensure legal compliance while using copyleft libraries as core dependencies in your software project.
Legal Reality: Whether linking creates a combined work is genuinely unsettled, and often has to be decided case by case. The FSF’s position is that linking a GPL library statically or dynamically creates a combined work; some legal scholars and Commission EUPL guidance disagree, particularly for dynamic linking through a stable API. Most Member States have no case law on this, so no firm general rule can be stated. The guidance below follows the conservative, widely-adopted reading.
JLA Selection Strategy: Under the conservative reading, linking to a GPL library means the combined program you distribute must be released under matching reciprocal terms, so configure JLA to require source code disclosure (
Disclose source) and reciprocal licensing (Copyleft/Share a.).
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. Copyright,Disclose source, andCopyleft/Share a.Support Column: Select
OSI approved
Example JLA Matches:
GPL-3.0,EUPL-1.2The safe default, not settled law: On the conservative reading, your own files must be under a GPL-compatible license (GPL itself, or permissive licenses such as MIT or BSD), and the combined program you distribute is under GPL.
EUPL-1.2also works for your own files through its compatibility clause, but the combined program then goes out under GPL anyway.What you ship matters: Static linking copies the library’s code into your binary, so you always distribute it. With dynamic linking (including a Python
import), the library stays a separate file. If you publish only your own source and users install the GPL library themselves, the risk is much lower, although the FSF would still expect your code to be GPL-compatible. If you bundle the library, in an executable, a container image, or a compiled binary, GPL clearly applies to what you ship.Alternatives if you want to stay permissive:
Find a permissively licensed alternative library.
Use an LGPL library instead: with dynamic linking, your own code can stay permissive, provided you keep its notices and do not restrict users from modifying the library or reverse engineering to debug those modifications.
Use an EUPL-1.2 library through dynamic linking: Commission guidance says this does not make your program a derivative work (guidance, not case law). Static linking or copying EUPL code is treated as a combined work.
Call a GPL tool as a separate program (e.g., via the command line) rather than importing it. This is generally treated as two programs communicating, not a combined work.
Copyleft licenses are not compatible with each other: Two strong copyleft licenses can each demand that the combined work use their terms, which makes the combination undistributable. The classic trap is
GPL-2.0-only: without the “or later” clause you cannot upgrade to GPL-3.0 to resolve a conflict, so GPL-2.0-only code cannot be combined with GPL-3.0 or Apache-2.0 code at all. Always check the exact SPDX identifierGPL-2.0-onlyandGPL-2.0-or-laterbehave very differently.Downstream Obligations: Anyone to whom you distribute the application must receive full access to your source code under GPL-compatible terms, along with upstream copyright notices and the build scripts needed to recompile it. Running the software internally, without distributing it, creates no such obligation though note that
AGPL-3.0extends this to network use, such as a web application built on an AGPL library.Allowed Inbound Code & Dependencies: Your project can import or include other permissively licensed packages (MIT, BSD, Apache-2.0) and public domain waivers (CC0). However, all code linked together in the final executable or runtime environment must satisfy GPL compatibility; for example,
Apache-2.0is compatible with GPL-3.0 but not with GPL-2.0.Check your dependencies: Tools such as
pip-licenses(Python) list the license of every installed package. Most package ecosystems have an equivalent. Run one once per project.In-File Identification (SPDX): Apply standard machine-readable SPDX tags directly at the top of your main scripts:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: GPL-3.0-or-later
import gpl_licensed_solver # External GPL dependency: conservative reading requires GPL compatibility
def solve_system(data):
return gpl_licensed_solver.compute(data)
Scenario 7: Distributing pre-built container images
You compiled and published a pre-built container image (e.g., pushing a compiled Docker image to Docker Hub, GitHub Container Registry, or an institutional registry, or sharing an Apptainer .sif file) containing an OS layer, runtime binaries, dependencies, and your application code.
Licensing Goal: Safely distribute compiled container images without violating the license terms of any software layer or binary included inside the image.
Legal Reality: A compiled container image is a multi-license aggregate bundle, not a single combined work. Distributing pre-built binaries makes you a distributor of every package inside, so source-availability obligations apply to the copyleft components (Linux base packages, coreutils, GPL libraries). But those packages sitting in the same filesystem as your application do not make your application a derivative of them: this is mere aggregation. Your own code keeps whatever license you chose; you simply also carry distributor obligations for the copyleft software you are shipping alongside it.
JLA Selection Strategy: Because a container image combines multiple distinct software components, JLA is used to evaluate constituent component obligations. When distributing compiled binaries containing copyleft layers, source disclosure requirements (
Disclose source) must be fulfilled for those specific layers.
Solution
What to select in the JLA interface:
Can Column: Select
DistributeandCommercial useMust Column: Select
Incl. CopyrightandDisclose sourceSupport Column: Select
OSI approved
JLA Outcome: No single license applies. Use JLA per component to check each one’s obligations, then record the aggregate in your image metadata.
Aggregation covers independent programs only: Mere aggregation applies to programs that simply live side by side in the image, such as your application next to
bashorcoreutils. If your application actually imports or links a GPL library inside the image, that relationship is a linking question, covered by Scenario 5.Multi-License Aggregation Nuance: Applying a permissive license (like MIT) to your application code inside the container does not override or erase the GPL/LGPL obligations of base system packages installed in
/usr/libor/usr/bin. Distributing the built image binary makes you a distributor of all installed packages.What counts as distribution: Pushing an image to a public registry, or sharing an image or
.siffile with people outside your organisation, is distribution. Keeping an image in a private registry used only within your own organisation is generally not. If you are unsure, treat it as distribution.Downstream Obligations: You must ensure downstream users can obtain the corresponding source for the copyleft components you shipped. Publishing your
Dockerfiledocuments the build but does not by itself satisfy this the GPL asks for the source of the binaries actually distributed. In practice, most research images rely on unmodified upstream distribution packages, and pointing to the distributor’s public source archives is common practice. The exact rules differ between GPL versions, however, and many distribution packages are GPL-2.0, so for images on public registries the safest option is to keep the relevant source available yourself. If you modify or rebuild a copyleft component yourself, you must provide that source directly.Watch for non-redistributable software: The bigger risk in an image is often proprietary software you are not allowed to redistribute at all, such as parts of NVIDIA CUDA, Intel’s math libraries, MATLAB runtimes, or commercial solvers. Their redistribution terms are set by each vendor’s license, so check them before publishing.
Generate a Software Bill of Materials (SBOM): Before publishing an image, use tools such as Syft or Trivy to list every package inside it, with versions and licenses. This SBOM is the complete record of what you distribute, and it shows whether any non-redistributable software is included. Consider publishing it alongside the image.
In-File Identification (Metadata Annotations): Document the multi-license nature of the aggregate bundle using standard OCI (Open Container Initiative) image labels inside your Dockerfile. The
licenseslabel is a summary: a base image contains many more licenses than it lists, so the SBOM remains the complete record.
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
FROM ubuntu:24.04
LABEL org.opencontainers.image.authors="author@institute.eu"
# OCI standard image annotations, displayed by registries and tools
LABEL org.opencontainers.image.title="My Research Pipeline"
# Summary only; see the published SBOM for the full list of licenses
LABEL org.opencontainers.image.licenses="MIT AND GPL-3.0-or-later"
LABEL org.opencontainers.image.vendor="My Institute Name"
LABEL org.opencontainers.image.description="Includes Ubuntu 24.04 base layers (GPL/LGPL) and custom solver (MIT)"
COPY solver.py /app/solver.py
Scenario 8: AI-assisted code generation
You used AI tools (e.g., GitHub Copilot, ChatGPT, Claude) to write functions, unit tests, or documentation for your research software repository.
Licensing Goal: Apply a permissive license (
MITorApache-2.0) to your repository with confidence, without incurring hidden copyright infringement or copyleft obligations from code the AI model reproduced from its training data.Legal Reality: Unmodified AI-generated outputs lack human authorship and are generally not eligible for copyright protection under current EU and international legal standards. Most real code, however, is a mix of human and AI contribution: you prompt, select, edit, and integrate. Where the line falls between AI output and your own work is unsettled and varies between Member States; there is no percentage or line-count threshold. The more you design, choose, edit, and integrate, the stronger your claim that the result is your work. Separately, if an LLM reproduces a substantial copyrighted code snippet verbatim from its training data (memorization), that output snippet retains its original copyright and license obligations.
JLA Selection Strategy: To ensure maximum adoption and academic reuse for your overall codebase, require that your copyright notice is kept (
Incl. Copyright) while avoiding share-alike constraints (leavingCopyleft/Share a.unselected), supported by automated compliance checks.
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. CopyrightSupport Column: Select
OSI approved
Example JLA Matches:
MIT,Apache-2.0,BSD-3-ClauseLicense your repository as normal: Your license covers everything you authored. Any purely AI-generated parts that are not protected by copyright are free to use anyway, so the license does no harm there. Many AI tools’ terms state that the output belongs to you as far as any rights exist, but a contract cannot create copyright that the law does not grant.
Checking for memorized code: Memorization is uncommon for everyday code, but it does happen, especially for well-known code that appears many times in training data. Practical checks:
Be most careful with long, distinctive functions and implementations of well-known algorithms. Short boilerplate and unit tests are low risk.
If you use GitHub Copilot, check whether the setting that blocks suggestions matching public code is enabled for your account or organisation.
If a suggestion looks suspiciously polished, search for a distinctive line of it on GitHub. If it appears in a copyleft project, treat it as that project’s code (Scenario 4).
Marking AI-generated code: Some projects and AI tool terms require contributors to disclose AI involvement via a commit trailer, a PR checkbox, or an in-file comment. Even where it is optional, marking AI-assisted sections is increasingly recommended practice: it records provenance, signals to reviewers where extra scrutiny is warranted, and makes later authorship or infringement questions much easier to resolve. Check the contribution guidelines of any project you submit to.
Use AI to write code, not to decide licensing: As noted in the section on the limitations of AI-assisted licensing advice, AI assistants tend to apply US legal concepts. Check licensing questions against the actual license text.
Downstream Obligations: Downstream users must preserve your copyright notice for the repository. They are free to reuse, modify, and integrate your code into commercial or open-source projects.
Allowed Inbound Snippets: You can include permissively licensed code, public domain code (CC0), and AI-generated snippets that you have checked for verbatim reproduction of training data, as described above.
In-File Identification (SPDX): Apply standard machine-readable SPDX tags directly at the top of your scripts, and mark AI-assisted code where it appears:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
def filter_sensor_data(raw_readings: list[float]) -> list[float]:
"""Cleans raw sensor data (written with AI assistance and human review)."""
return [reading for reading in raw_readings if reading > 0.0]
Scenario 9: Packaging AI workflows, datasets, and model weights
You are developing research software that includes source code alongside trained machine learning model weights (.pt, .safetensors) and benchmark datasets.
Licensing Goal: Apply a clear licensing structure, with different licenses for different parts of the repository, that makes both the software source code and the non-code assets (data, weights) open and reusable under appropriate legal frameworks.
Legal Reality: Standard software licenses (MIT, GPL) are written for source code and fit datasets and model parameters poorly. Datasets may attract the EU sui generis database right where there has been substantial investment in obtaining, verifying, or presenting their contents. Model weights are a harder case: they are the numerical values learned during training, neither code nor a database, and whether they attract any copyright protection in the EU is genuinely unsettled. Because of this uncertainty, applying an explicit license to weights is about setting clear terms for your users, not about relying on a settled legal right.
JLA Selection Strategy: Use JLA to select an OSI-approved open-source license for the executable code component (
Incl. Copyrightselected), while using Creative Commons licenses (e.g.,CC-BY-4.0orCC0-1.0) for the dataset and weight files.
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. CopyrightSupport Column: Select
OSI approved
Example JLA Matches:
MIT,Apache-2.0(for the code component)Code vs. Data/Weights: Avoid applying software licenses like GPL or MIT to raw datasets or model weights their terms reference source code, object code, and linking, which leaves users guessing about what applies. Use CC-BY-4.0 or CC0-1.0 for non-code assets instead. CC-BY-4.0 requires credit; CC0-1.0 requires nothing, which makes it easier for data that others will combine with many other datasets. Note that this pattern is sometimes called “dual-licensing”, but that term usually means offering the same work under two licenses.
You can only license what is yours: If your dataset contains material you did not create, such as scraped text, images, or other people’s data, your license covers only your own contribution; the original content keeps its own rights. If your dataset contains personal data, data protection rules apply regardless of the license.
Open weights are not always open source: If you fine-tuned an existing model, its license still applies to what you built on it. Many models published with open weights come with their own licenses restricting, for example, commercial use or certain applications. Check the base model’s license before fine-tuning and publishing.
Behavioral licenses (OpenRAIL): Licenses such as OpenRAIL impose usage restrictions (e.g., prohibiting specific harmful uses). This can be a reasonable choice, but it means they do not qualify as OSI-approved open source and will not appear in standard JLA queries.
Downstream Obligations: Downstream users must keep your copyright notice and license text for the code (under your chosen software license) and give credit for the model weights and data as the corresponding Creative Commons license requires. For academic citation, add a
CITATION.cfffile to your repository.Allowed Inbound Assets: You may combine permissively licensed Python code with CC-BY-4.0 datasets, provided the attribution files clearly separate code licenses from data/weight licenses. Models published with open weights can be included only under the terms of their own licenses (see above).
In-File Identification (SPDX / Licensing Structure): Binary files such as weights and datasets cannot contain comments, so REUSE marks them with a companion
.licensefile next to each one (e.g.,climate_weights.safetensors.license) or with a singleREUSE.tomlfile covering whole folders. State in your README which license covers which folder, and fill in the license fields on platforms such as Zenodo or Hugging Face. In your code, document the structure in the header:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
#
# Note: Source code is licensed under MIT.
# Model weights in /models/ and datasets in /data/ are licensed under CC-BY-4.0.
from safetensors.torch import load_file
def load_pipeline():
weights = load_file("models/climate_weights.safetensors")
return weights
Software citation
Discussion (Citation-1): Explain how you currently cite software
Do you cite software that you use? How?
If I wanted to cite your code/scripts, what would I need to do?
Social coding
In social-coding.md:
Social-1: Think about if and how you share
Did you ever share your code? If yes, what motivated you? Come up with reasons for sharing your scripts/code/data.
Also think about reasons for not sharing.
In social-coding.md:
Social-2: Discussion about “You aren’t required to support anyone”
Have you experienced an implicit expectation of support?
Supporting all requests can lead to overworking and mental health issues.
Not supporting requests can also induce guilt.
Most projects are maintained by 1 or 2 persons.
Most projects cannot retain contributors for a longer time. Interests change. “Casual contributors are like tourists visiting NYC for a weekend” (Nadia Asparouhova, book below).
If you maintain all projects that you start forever, at some point it may be difficult to start new projects.
What are your experiences? Do you agree with the above thoughts?
Book recommendation: Nadia Asparouhova (formerly Nadia Eghbal): “Working in Public: The Making and Maintenance of Open Source Software (Stripe Press)”