Software licensing focusing on open source
Objectives
Principles of open source licensing
Difference between permissive and copyleft licenses
Frameworks for AI-generated and AI-assisted code
Determine the software license for your project following EU copyright framework
Navigate the Joinup Licensing Assistant to select a compliant license
Understand the licensing distinction between container recipes and container images
Limitations and context of this lesson
This lesson is designed as practical educational material for researchers and research software engineers, not formal legal advice.
EU directives set only minimum requirements in some areas: Member States implement them differently and may add national rules not covered here. For example, some Member States let university researchers retain ownership of the programs they write instead of applying the employer rule in Art. 2(3).
Institutional Context: Employment contracts, grant agreements, and university policies heavily influence software ownership and licensing choices.
This lesson covers only the general principles of open-source reuse, copyright scope, and software adaptation.
If you need formal guidance, the references below can help and so can legal experts, especially if your host institute has a legal services office:
Introduction: What is a Software License?
Under copyright law worldwide, software without an explicit license defaults to All Rights Reserved: nobody else may run, copy, modify, distribute, or build on your code. A software license is how the copyright holder exercises their exclusive rights, granting others permission to reproduce, distribute, modify, and sometimes sublicense the work.
Note that author and copyright holder may differ: under Art. 2(3), an employer exercises the economic rights in code written by an employee on the job, unless a contract says otherwise. The employee is still the author; the employer is who licenses it. This matters in practice, because the person choosing the license for a research project is often not the person who wrote the code.
This lesson focuses on open-source licenses. If your employment terms and institutional policy allow you to open-source the code you write, we recommend doing so. It makes you a better citizen of the research community, since others can reuse, verify, and build on your work. It also protects your future self: code your employer owns and never licenses stays locked behind All Rights Reserved when you change jobs, whereas an open license grants everyone the right to reuse it, including you.
Open-source licenses fall into three families, which differ in what they let downstream users do:
Three rules of thumb as an compliment to the diagram:
Copyleft only applies when you share the code. Running modified GPL code on your own machine or cluster creates no obligations.
“Weak” copyleft still has conditions. For example, if you ship an LGPL library inside a closed product, you must still let users modify and debug that library.
Copyleft licenses often don’t mix. Code under two different copyleft licenses may not be combinable, so your choice today decides who can build on your work later.
You will hear copyleft called viral or infectious. The slang is misleading: copyleft doesn’t spread just because GPL code sits next to yours in a repository or container. It only applies when you build GPL code into your own, for example by copying in a snippet. And choosing copyleft is a legitimate project decision, not a sign that a license is harmful.
This lesson covers both directions: choosing terms for software you write, and complying with terms attached to code written by others. The scenarios later work through each case.
Copyright Foundation: Expression vs. Ideas
Under Directive 2009/24/EC, software is protected by copyright as a literary work (Art. 1(1)). But copyright protects only the expression, not the ideas beneath it: Art. 1(2) explicitly excludes “ideas and principles which underlie any element of a computer program, including those which underlie its interfaces.”
Protected: your specific source code text, binaries, container recipes, prompt text, and preparatory design material.
Not protected: mathematical algorithms, scientific models, programming logic, data structures, and interfaces.
The CJEU confirmed this line in SAS Institute v World Programming (C-406/10): a program’s functionality, its programming language, and its data file formats are ideas, not expression, and are therefore outside copyright. Someone may reimplement your algorithm from scratch; they may not copy your code. This is exactly why licenses exist they set the terms for the expression, which is the only part copyright lets you control.
Plagiarism vs. Intellectual Property Rights = Research Ethics vs. Law
This insert can be skipped and left as reading exercise
In academic context it is important to consider also plagiarism and how it relates to copyright and more broadly Intellectual Property Rights (a clear explanation at this page). Plagiarism is the practice of taking somebody else’s ideas or work and claim them as your own: it is the unacknowledged use of another person’s work. Intellectual Property Rights (IPRs) infringement instead is the unauthorised use of another’s work.
IPRs can be classified in two main groups (WTO): i) Copyright and rights related to copyright (computer programs are here) and ii) Industrial properties like trademarks, and inventions (which may include specific technical implementations of systems or code) protected by patents.
In research ethics, plagiarism is one of the three main forms of research misconduct (along with fabrication and falsification, see ALLEA, European Code of Conduct for Research Integrity). Plagiarism is not illegal per se, but it can lead to serious consequences like the retraction of published work. One can engage in plagiarism, without necessarily breaking any IPR law (e.g. write a new book by reusing the plot of an old book that is not under copyright anymore). Copyright infringment instead is illegal and it can result in criminal charges (e.g. fines). Copyright however protects the particular expression of an idea or fact (for example, the specific source code of a program, but not the underlying algorithm itself).
There is no pre-defined “number of lines of code”, “seconds of a song”, or “pixels of an image” that can clearly set the basis for plagiarism or IPR infringement. However in the context of research, it can be possible to use Quotation Exception (in EU, ref) and Fair use (in USA, ref). Fair use has become controversial recently as it is used as legal basis for training large language models based on scraped internet data (See for example Henderson, P., Li, X., Jurafsky, D., Hashimoto, T., Lemley, M. A., & Liang, P. (2023). Foundation models and fair use. Journal of Machine Learning Research, 24(400), 1-79.)
Scope of this Lesson: What Counts as Software?
Across international frameworks (17 U.S.C. § 101 and WIPO model provisions), software is broadly defined as a set of statements or instructions used directly or indirectly in a computer to bring about a certain result. Research software goes well beyond Python scripts, so this lesson covers six asset types find the ones matching your own project, since the scenarios later map onto them:
Source Code original algorithms, or implementations of published methods.
Third-Party Integrations embedded snippets and linked libraries (static or dynamic).
Infrastructure as Code Ansible playbooks, Terraform configs, container recipes (
Dockerfile, Apptainer.def).Container Images built binary snapshots (
.siffiles, OCI registry images).AI-Assisted Code generated or refactored with human oversight.
AI Prompt Templates engineered system prompts meeting the threshold of human authorship.
Motivation: Debugging a License Compliance Failure
With the three license families in mind, examine what happens when they collide inside an automated CI/CD pipeline:
Option D : deleting the
LICENSEfile makes the scanner quiet without changing anything legally. You are still distributing someone else’s copyleft code without honouring its terms, and you have now stripped your own users of any permission to use your work. A green pipeline is not a compliance result.Option C works only if you genuinely reimplement the functionality without copying the original expression. As the idea/expression split above establishes, the algorithm is free to reuse the specific code is not. Reading the original closely and retyping a close paraphrase is still copying.
Limitations of AI-Assisted Licensing Advice
Modern software developers and RSEs routinely rely on AI coding assistants (ChatGPT, Claude, GitHub Copilot) to generate boilerplate, refactor functions, and answer project setup questions.
However, using these tools for legal or licensing guidance introduces a subtle risk of AI legal bias. AI models are overwhelmingly trained on US-centric web data and legal forum posts, so their outputs default almost universally to US common law concepts such as Fair Use, Work Made for Hire, and Derivative Works.
Developers working under EU statutory frameworks face a different legal reality around exceptions, ownership, and code adaptation. The clearest example is the term you will hear constantly:
US law (17 U.S.C. § 101) formally defines “derivative work”, and AI assistants reach for it to describe almost any code modification.
EU law (Directive 2009/24/EC, Art. 4(1)(b)) does not use that term at all. It grants exclusive rights over “the translation, adaptation, arrangement and any other alteration of a computer program” governed collectively as an adaptation.
Licenses vary:
EUPL-1.2defines “Derivative Works” in its own text as a contractual term, andGPL-2.0used the phrase too.GPL-3.0deliberately dropped it in favour of “modify” and “a work based on the Program”, because its drafters recognised the term means different things in different jurisdictions the same problem you face when an AI assistant uses it.
So when an AI assistant tells you a snippet creates a “derivative work”, treat that as a prompt to check the actual question under EU law: is this a statutory adaptation, or a combined work across a technical boundary? The rest of this lesson gives you that EU-aligned framework.
Standardizing In-File Declarations: SPDX Identifiers
Selecting a license is only half the job. Automated scanners and CI/CD pipelines need a machine-readable way to verify compliance per file without parsing legal text.
Managed by the Linux Foundation, SPDX provides standardized short identifiers for licenses (MIT, Apache-2.0, GPL-3.0-only, EUPL-1.2). The REUSE specification, maintained by the FSFE, builds on SPDX to define how every file should declare its copyright and license. Each file starts with two tags:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
SPDX-FileCopyrightTextnames the copyright holder and year. This may be your institution rather than you personally; check your institution’s policy.SPDX-License-Identifiernames the license, using the exact SPDX identifier. Watch the suffix:GPL-3.0-onlyandGPL-3.0-or-laterbehave differently, so choose deliberately.
The tags only point to a license, so the full license text must also be in your repository. REUSE places one text file per license in a LICENSES/ folder (e.g., LICENSES/MIT.txt). Running reuse lint then checks that every file carries both tags and that every license it names has its text present.
Every scenario below shows the SPDX tagging for its asset type Python scripts, container recipes, and prompt templates each have their own conventions.
License Selection Decision Matrix & Scenario Index
Our decision framework is grounded in the European Commission’s Joinup Licensing Assistant (JLA), which sorts licenses across six criteria: Can (permissions), Must (obligations), Cannot (restrictions), Compatible (interoperability), Law (jurisdiction), and Support (governance).
The scenarios below are independent. Find the row that matches what you are actually building, jump to it, and skip the rest.
If you are… |
Scenario |
Typical Outcome |
Example Licenses |
|---|---|---|---|
Writing everything yourself |
🟢 Free choice |
|
|
Implementing a published algorithm |
🟢 Free choice copyleft if you want reciprocity |
|
|
Pasting in a permissive snippet |
🟢 Stay permissive, keep notices |
|
|
Pasting in a copyleft snippet |
🟡 Strong copyleft likely required |
|
|
Importing or linking a library |
🟡 Depends on which copyleft see below |
|
|
Writing a Dockerfile or |
🟢 Free choice |
|
|
Publishing a built image |
⚠️ Multi-license bundle |
Governed by each layer’s own terms |
|
Using Copilot, ChatGPT or Claude |
🟢 Free choice, verify for memorization |
|
|
Shipping prompts, weights or datasets |
🟢 Dual-license code vs. assets |
|
This lesson covers nine scenarios; a typical session works through three or four. The rest are here for reference when your project changes.
Exercise-1: How do you work with others’ software and ideas?
Licensing-1: Which scenarios typically describe your work?
The text below can be copied to the collaborative document for an online poll:
## Question: How do you work with other's code?
**Choose many**. Vote by adding an `o` character:
- 1. Writing everything yourself
- votes:
- 2. Implementing a published algorithm
- votes:
- 3. Pasting in a permissive snippet
- votes:
- 4. Pasting in a copyleft snippet
- votes:
- 5. Importing or linking a library
- votes:
- 6. Writing a Dockerfile or .def
- votes:
- 7. Publishing a built image
- votes:
- 8. Using Copilot, ChatGPT, Claude, or similar
- votes:
- 9. Shipping prompts, weights or datasets
- votes:
Exercise-2: How small is a snippet?
Licensing-2: Can you decide from the number of lines?
The text below can be copied to the collaborative document for an online poll:
## Question: Which of these can you safely copy *based only on its size*?
**Choose many**. Vote by adding an `o` character:
- A. A one-line expression: `return max(lo, min(x, hi))`
- votes:
- B. Five lines of ordinary boilerplate for parsing command-line arguments
- votes:
- C. Three unusually written lines copied verbatim from a GPL-licensed solver
- votes:
- D. Twenty lines you wrote independently after reading an algorithm in a paper, without looking at another implementation
- votes:
- E. Anything under 10 lines is too small to be copyrighted
- votes:
- F. None of the above: the number of lines alone does not decide
- votes:
### Follow-up question
For each example, what information would you want to know before reusing or publishing the code?
Solution
The key answer is F: there is no fixed safe number of lines.
Copyright does not use a numerical threshold such as 5, 10, or 20 lines. The important question is whether what has been copied is protected expression. Very short or purely functional code may not meet the originality threshold, while a short but distinctive piece of code may.
A and B: they may be too simple, conventional, or constrained by function to contain protectable expression, but their size alone does not answer the question.
C: being only three lines does not automatically make copied code unprotected. Check its provenance and license.
D: independently implementing the idea or algorithm is different from copying somebody else’s expression of it.
E: there is no “10-line rule”.
Practical rule: if you copied code and are unsure whether it is protected, check where it came from and under which license it was published. Preserve any required notices, or independently implement the underlying idea instead of copying the code.
Module 2: The Dependency Minefield – Inbound Code & Linking
Embedding third-party snippets or linking against external libraries introduces boundaries that can constrain your license choice. How far those boundaries reach depends on which license the inbound code carries.
Module 3: Dependency Linking & Packaging
When software incorporates external dependencies, whether by dynamic linking, static compiling, or bundling binaries into container images licensing obligations expand beyond your own written source code. This module covers how dependency boundaries, build automation scripts, and packaged container artifacts affect legal compliance under the Joinup Licensing Assistant (JLA) framework.
Scenario 5: Linking against a GPL-licensed library
You are developing a software application that imports or links against an external software library licensed under GPL-3.0 (e.g., importing a GPL Python package or linking a C/C++ static/shared library).
Licensing Goal: Ensure legal compliance while using copyleft libraries as core dependencies in your software project.
Legal Reality: Whether linking creates a combined work is genuinely unsettled, and often has to be decided case by case. The FSF’s position is that linking a GPL library statically or dynamically creates a combined work; some legal scholars and Commission EUPL guidance disagree, particularly for dynamic linking through a stable API. Most Member States have no case law on this, so no firm general rule can be stated. The guidance below follows the conservative, widely-adopted reading.
JLA Selection Strategy: Under the conservative reading, linking to a GPL library means the combined program you distribute must be released under matching reciprocal terms, so configure JLA to require source code disclosure (
Disclose source) and reciprocal licensing (Copyleft/Share a.).
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. Copyright,Disclose source, andCopyleft/Share a.Support Column: Select
OSI approved
Example JLA Matches:
GPL-3.0,EUPL-1.2The safe default, not settled law: On the conservative reading, your own files must be under a GPL-compatible license (GPL itself, or permissive licenses such as MIT or BSD), and the combined program you distribute is under GPL.
EUPL-1.2also works for your own files through its compatibility clause, but the combined program then goes out under GPL anyway.What you ship matters: Static linking copies the library’s code into your binary, so you always distribute it. With dynamic linking (including a Python
import), the library stays a separate file. If you publish only your own source and users install the GPL library themselves, the risk is much lower, although the FSF would still expect your code to be GPL-compatible. If you bundle the library, in an executable, a container image, or a compiled binary, GPL clearly applies to what you ship.Alternatives if you want to stay permissive:
Find a permissively licensed alternative library.
Use an LGPL library instead: with dynamic linking, your own code can stay permissive, provided you keep its notices and do not restrict users from modifying the library or reverse engineering to debug those modifications.
Use an EUPL-1.2 library through dynamic linking: Commission guidance says this does not make your program a derivative work (guidance, not case law). Static linking or copying EUPL code is treated as a combined work.
Call a GPL tool as a separate program (e.g., via the command line) rather than importing it. This is generally treated as two programs communicating, not a combined work.
Copyleft licenses are not compatible with each other: Two strong copyleft licenses can each demand that the combined work use their terms, which makes the combination undistributable. The classic trap is
GPL-2.0-only: without the “or later” clause you cannot upgrade to GPL-3.0 to resolve a conflict, so GPL-2.0-only code cannot be combined with GPL-3.0 or Apache-2.0 code at all. Always check the exact SPDX identifierGPL-2.0-onlyandGPL-2.0-or-laterbehave very differently.Downstream Obligations: Anyone to whom you distribute the application must receive full access to your source code under GPL-compatible terms, along with upstream copyright notices and the build scripts needed to recompile it. Running the software internally, without distributing it, creates no such obligation though note that
AGPL-3.0extends this to network use, such as a web application built on an AGPL library.Allowed Inbound Code & Dependencies: Your project can import or include other permissively licensed packages (MIT, BSD, Apache-2.0) and public domain waivers (CC0). However, all code linked together in the final executable or runtime environment must satisfy GPL compatibility; for example,
Apache-2.0is compatible with GPL-3.0 but not with GPL-2.0.Check your dependencies: Tools such as
pip-licenses(Python) list the license of every installed package. Most package ecosystems have an equivalent. Run one once per project.In-File Identification (SPDX): Apply standard machine-readable SPDX tags directly at the top of your main scripts:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: GPL-3.0-or-later
import gpl_licensed_solver # External GPL dependency: conservative reading requires GPL compatibility
def solve_system(data):
return gpl_licensed_solver.compute(data)
Scenario 7: Distributing pre-built container images
You compiled and published a pre-built container image (e.g., pushing a compiled Docker image to Docker Hub, GitHub Container Registry, or an institutional registry, or sharing an Apptainer .sif file) containing an OS layer, runtime binaries, dependencies, and your application code.
Licensing Goal: Safely distribute compiled container images without violating the license terms of any software layer or binary included inside the image.
Legal Reality: A compiled container image is a multi-license aggregate bundle, not a single combined work. Distributing pre-built binaries makes you a distributor of every package inside, so source-availability obligations apply to the copyleft components (Linux base packages, coreutils, GPL libraries). But those packages sitting in the same filesystem as your application do not make your application a derivative of them: this is mere aggregation. Your own code keeps whatever license you chose; you simply also carry distributor obligations for the copyleft software you are shipping alongside it.
JLA Selection Strategy: Because a container image combines multiple distinct software components, JLA is used to evaluate constituent component obligations. When distributing compiled binaries containing copyleft layers, source disclosure requirements (
Disclose source) must be fulfilled for those specific layers.
Solution
What to select in the JLA interface:
Can Column: Select
DistributeandCommercial useMust Column: Select
Incl. CopyrightandDisclose sourceSupport Column: Select
OSI approved
JLA Outcome: No single license applies. Use JLA per component to check each one’s obligations, then record the aggregate in your image metadata.
Aggregation covers independent programs only: Mere aggregation applies to programs that simply live side by side in the image, such as your application next to
bashorcoreutils. If your application actually imports or links a GPL library inside the image, that relationship is a linking question, covered by Scenario 5.Multi-License Aggregation Nuance: Applying a permissive license (like MIT) to your application code inside the container does not override or erase the GPL/LGPL obligations of base system packages installed in
/usr/libor/usr/bin. Distributing the built image binary makes you a distributor of all installed packages.What counts as distribution: Pushing an image to a public registry, or sharing an image or
.siffile with people outside your organisation, is distribution. Keeping an image in a private registry used only within your own organisation is generally not. If you are unsure, treat it as distribution.Downstream Obligations: You must ensure downstream users can obtain the corresponding source for the copyleft components you shipped. Publishing your
Dockerfiledocuments the build but does not by itself satisfy this the GPL asks for the source of the binaries actually distributed. In practice, most research images rely on unmodified upstream distribution packages, and pointing to the distributor’s public source archives is common practice. The exact rules differ between GPL versions, however, and many distribution packages are GPL-2.0, so for images on public registries the safest option is to keep the relevant source available yourself. If you modify or rebuild a copyleft component yourself, you must provide that source directly.Watch for non-redistributable software: The bigger risk in an image is often proprietary software you are not allowed to redistribute at all, such as parts of NVIDIA CUDA, Intel’s math libraries, MATLAB runtimes, or commercial solvers. Their redistribution terms are set by each vendor’s license, so check them before publishing.
Generate a Software Bill of Materials (SBOM): Before publishing an image, use tools such as Syft or Trivy to list every package inside it, with versions and licenses. This SBOM is the complete record of what you distribute, and it shows whether any non-redistributable software is included. Consider publishing it alongside the image.
In-File Identification (Metadata Annotations): Document the multi-license nature of the aggregate bundle using standard OCI (Open Container Initiative) image labels inside your Dockerfile. The
licenseslabel is a summary: a base image contains many more licenses than it lists, so the SBOM remains the complete record.
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
FROM ubuntu:24.04
LABEL org.opencontainers.image.authors="author@institute.eu"
# OCI standard image annotations, displayed by registries and tools
LABEL org.opencontainers.image.title="My Research Pipeline"
# Summary only; see the published SBOM for the full list of licenses
LABEL org.opencontainers.image.licenses="MIT AND GPL-3.0-or-later"
LABEL org.opencontainers.image.vendor="My Institute Name"
LABEL org.opencontainers.image.description="Includes Ubuntu 24.04 base layers (GPL/LGPL) and custom solver (MIT)"
COPY solver.py /app/solver.py
Module 4: Emerging Workflows & AI
AI-assisted development tools and machine learning models introduce unique legal challenges regarding copyright ownership, training data memorization, and behavioral restrictions. This module addresses how to license projects built with AI code generation tools and how to package research software that bundles AI models, weights, and datasets alongside source code.
Scenario 8: AI-assisted code generation
You used AI tools (e.g., GitHub Copilot, ChatGPT, Claude) to write functions, unit tests, or documentation for your research software repository.
Licensing Goal: Apply a permissive license (
MITorApache-2.0) to your repository with confidence, without incurring hidden copyright infringement or copyleft obligations from code the AI model reproduced from its training data.Legal Reality: Unmodified AI-generated outputs lack human authorship and are generally not eligible for copyright protection under current EU and international legal standards. Most real code, however, is a mix of human and AI contribution: you prompt, select, edit, and integrate. Where the line falls between AI output and your own work is unsettled and varies between Member States; there is no percentage or line-count threshold. The more you design, choose, edit, and integrate, the stronger your claim that the result is your work. Separately, if an LLM reproduces a substantial copyrighted code snippet verbatim from its training data (memorization), that output snippet retains its original copyright and license obligations.
JLA Selection Strategy: To ensure maximum adoption and academic reuse for your overall codebase, require that your copyright notice is kept (
Incl. Copyright) while avoiding share-alike constraints (leavingCopyleft/Share a.unselected), supported by automated compliance checks.
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. CopyrightSupport Column: Select
OSI approved
Example JLA Matches:
MIT,Apache-2.0,BSD-3-ClauseLicense your repository as normal: Your license covers everything you authored. Any purely AI-generated parts that are not protected by copyright are free to use anyway, so the license does no harm there. Many AI tools’ terms state that the output belongs to you as far as any rights exist, but a contract cannot create copyright that the law does not grant.
Checking for memorized code: Memorization is uncommon for everyday code, but it does happen, especially for well-known code that appears many times in training data. Practical checks:
Be most careful with long, distinctive functions and implementations of well-known algorithms. Short boilerplate and unit tests are low risk.
If you use GitHub Copilot, check whether the setting that blocks suggestions matching public code is enabled for your account or organisation.
If a suggestion looks suspiciously polished, search for a distinctive line of it on GitHub. If it appears in a copyleft project, treat it as that project’s code (Scenario 4).
Marking AI-generated code: Some projects and AI tool terms require contributors to disclose AI involvement via a commit trailer, a PR checkbox, or an in-file comment. Even where it is optional, marking AI-assisted sections is increasingly recommended practice: it records provenance, signals to reviewers where extra scrutiny is warranted, and makes later authorship or infringement questions much easier to resolve. Check the contribution guidelines of any project you submit to.
Use AI to write code, not to decide licensing: As noted in the section on the limitations of AI-assisted licensing advice, AI assistants tend to apply US legal concepts. Check licensing questions against the actual license text.
Downstream Obligations: Downstream users must preserve your copyright notice for the repository. They are free to reuse, modify, and integrate your code into commercial or open-source projects.
Allowed Inbound Snippets: You can include permissively licensed code, public domain code (CC0), and AI-generated snippets that you have checked for verbatim reproduction of training data, as described above.
In-File Identification (SPDX): Apply standard machine-readable SPDX tags directly at the top of your scripts, and mark AI-assisted code where it appears:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
def filter_sensor_data(raw_readings: list[float]) -> list[float]:
"""Cleans raw sensor data (written with AI assistance and human review)."""
return [reading for reading in raw_readings if reading > 0.0]
Scenario 9: Packaging AI workflows, datasets, and model weights
You are developing research software that includes source code alongside trained machine learning model weights (.pt, .safetensors) and benchmark datasets.
Licensing Goal: Apply a clear licensing structure, with different licenses for different parts of the repository, that makes both the software source code and the non-code assets (data, weights) open and reusable under appropriate legal frameworks.
Legal Reality: Standard software licenses (MIT, GPL) are written for source code and fit datasets and model parameters poorly. Datasets may attract the EU sui generis database right where there has been substantial investment in obtaining, verifying, or presenting their contents. Model weights are a harder case: they are the numerical values learned during training, neither code nor a database, and whether they attract any copyright protection in the EU is genuinely unsettled. Because of this uncertainty, applying an explicit license to weights is about setting clear terms for your users, not about relying on a settled legal right.
JLA Selection Strategy: Use JLA to select an OSI-approved open-source license for the executable code component (
Incl. Copyrightselected), while using Creative Commons licenses (e.g.,CC-BY-4.0orCC0-1.0) for the dataset and weight files.
Solution
What to select in the JLA interface:
Can Column: Select
Distribute,Modify/merge, andCommercial useMust Column: Select
Incl. CopyrightSupport Column: Select
OSI approved
Example JLA Matches:
MIT,Apache-2.0(for the code component)Code vs. Data/Weights: Avoid applying software licenses like GPL or MIT to raw datasets or model weights their terms reference source code, object code, and linking, which leaves users guessing about what applies. Use CC-BY-4.0 or CC0-1.0 for non-code assets instead. CC-BY-4.0 requires credit; CC0-1.0 requires nothing, which makes it easier for data that others will combine with many other datasets. Note that this pattern is sometimes called “dual-licensing”, but that term usually means offering the same work under two licenses.
You can only license what is yours: If your dataset contains material you did not create, such as scraped text, images, or other people’s data, your license covers only your own contribution; the original content keeps its own rights. If your dataset contains personal data, data protection rules apply regardless of the license.
Open weights are not always open source: If you fine-tuned an existing model, its license still applies to what you built on it. Many models published with open weights come with their own licenses restricting, for example, commercial use or certain applications. Check the base model’s license before fine-tuning and publishing.
Behavioral licenses (OpenRAIL): Licenses such as OpenRAIL impose usage restrictions (e.g., prohibiting specific harmful uses). This can be a reasonable choice, but it means they do not qualify as OSI-approved open source and will not appear in standard JLA queries.
Downstream Obligations: Downstream users must keep your copyright notice and license text for the code (under your chosen software license) and give credit for the model weights and data as the corresponding Creative Commons license requires. For academic citation, add a
CITATION.cfffile to your repository.Allowed Inbound Assets: You may combine permissively licensed Python code with CC-BY-4.0 datasets, provided the attribution files clearly separate code licenses from data/weight licenses. Models published with open weights can be included only under the terms of their own licenses (see above).
In-File Identification (SPDX / Licensing Structure): Binary files such as weights and datasets cannot contain comments, so REUSE marks them with a companion
.licensefile next to each one (e.g.,climate_weights.safetensors.license) or with a singleREUSE.tomlfile covering whole folders. State in your README which license covers which folder, and fill in the license fields on platforms such as Zenodo or Hugging Face. In your code, document the structure in the header:
# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
#
# Note: Source code is licensed under MIT.
# Model weights in /models/ and datasets in /data/ are licensed under CC-BY-4.0.
from safetensors.torch import load_file
def load_pipeline():
weights = load_file("models/climate_weights.safetensors")
return weights
Best Practices: Attaching a License to Your Repository
Once you have selected a license using the JLA, you must officially attach it to your repository so automated scanners, package registries, and downstream users can verify your terms.
1. Adding the Root LICENSE File
Always place the full text of your chosen license in a plain-text file named LICENSE or LICENSE.txt at the root of your repository.
Exact Legal Text: Copy the standard text directly from spdx.org/licenses or choosealicense.com.
Copyright Header: Ensure you fill in the copyright year and copyright holder line at the top of the license text:
Copyright (c) 2026 [Author Name or Institution Name]
Do Not Edit Terms: Never modify the legal wording of standard licenses (e.g., removing clauses from GPL or MIT). Edited texts are no longer the license they claim to be: compliance scanners cannot classify them, package registries may flag them, and downstream users have to get their own legal review before touching your code. If a standard license does not fit, pick a different standard license.
2. Documenting License Status in README.md
Add a dedicated License section near the bottom of your repository’s README.md file, along with a machine-readable badge:
## License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
[](https://opensource.org/licenses/MIT)
3. Automated Compliance with the REUSE Standard
For multi-asset research repositories containing code, data, container build recipes, and prompt templates, follow the FSFE REUSE Initiative standard:
Include License Texts: Place full license files inside a
LICENSES/directory (e.g.,LICENSES/MIT.txt,LICENSES/GPL-3.0-or-later.txt).Add In-File SPDX Headers: Label every source file, build script, and prompt file with SPDX tags.
Verify Compliance: Run the automated REUSE linter in your CI/CD pipeline:
# Install and run REUSE compliance check
pip install reuse
reuse lint
When reuse lint passes, every asset in your codebase carries a declared, machine-readable license that downstream users can check.
Summary: Resolving the Compliance Pipeline
When developing research software, license compliance is not an afterthought to debug at the end of a project, it is a proactive design choice. By using the Joinup Licensing Assistant (JLA) framework to align your repository license with your inbound dependencies from day one, your pipeline is far less likely to fail on a license conflict late in the project.
The diagram below illustrates how selecting a compatible license upfront ensures your code passes automated compliance checks and results in a legally sound release:
Compare this with the failing pipeline at the start of the lesson: the pipeline itself is identical. Nothing about the scanner changed the only difference is two decisions made before pushing.
Scenario Mapping Across the Pipeline
Choosing Your Own Terms (Scenario 1, Scenario 2 & Scenario 3): When you write original code, implement a published algorithm, or embed only permissive snippets, no inbound license constrains you the choice follows your goal. Pick permissive (
MIT,Apache-2.0) for maximum adoption, or copyleft (EUPL-1.2,GPL-3.0) if you want downstream improvements shared back. Either way, preserve any third-party notices attached to code you embedded.Handling Inbound Copyleft (Scenario 4 & Scenario 5): Copying a non-trivial copyleft snippet (e.g., CC BY-SA code from Stack Overflow, or a GPL fragment) creates a combined work. Linking against a copyleft library may do the same, depending on the license and the linking method. In both cases, selecting a compatible copyleft license upfront (
GPL-3.0orEUPL-1.2) satisfies the reciprocal terms and lets the scanner pass and checking the exact SPDX identifier first avoids theGPL-2.0-onlyincompatibility trap.Packaging and Build Automation (Scenario 6 & Scenario 7): Keep plain-text build recipes (Dockerfiles) permissively licensed for maximum reuse, while annotating compiled container image binaries as multi-license aggregate bundles to satisfy embedded base-layer obligations.
AI Assets and Dual-Licensing (Scenario 8 & Scenario 9): Run code-similarity scanners to catch LLM training memorization before releasing AI-assisted code, and apply dual-licensing to separate executable software code (
MIT) from non-code datasets and model weights (CC-BY-4.0).Standardized Distribution: Adding machine-readable SPDX headers across every script, Dockerfile, and prompt template lets
reuse lintconfirm that every asset has a declared, documented license. Note what this does and does not prove: the linter verifies that declarations exist and are well-formed, not that they are legally correct or mutually compatible. Automation makes your intent auditable it does not replace the judgment calls in the scenarios above.
Glossary
Glossary of terms (click to expand)
- 0BSD
Zero-Clause BSD, a permissive license so minimal it doesn’t even require keeping the copyright notice.
- Adaptation
EU term (Art. 4(1)(b)) for translating, arranging, or altering a program; roughly the US derivative work.
- AGPL-3.0
Strong copyleft that also requires sharing source when users interact with modified software over a network.
- All Rights Reserved
Default for unlicensed software: nobody but the copyright holder may run, copy, modify, or share it.
- Apache-2.0
Permissive license like MIT, plus an explicit patent grant and a requirement to note changes you made.
- API
Application Programming Interface: the defined way one program calls another. The idea of an interface is not protected by copyright; the code implementing it is.
- Author
The person who created the program; not always the copyright holder.
- BSD-3-Clause
Permissive license like MIT, plus a clause forbidding use of the authors’ names to promote derived products.
- CC BY-SA
Creative Commons share-alike license used for code posted on Stack Overflow; adaptations must carry the same terms, much like copyleft.
- CC-BY-4.0
Creative Commons license allowing any reuse if the creator is credited; suited to data, documentation, and models rather than code.
- CC0
Creative Commons tool waiving rights as far as the law allows, placing a work as close to the public domain as possible.
- CI/CD
Continuous Integration / Continuous Delivery: automated pipelines that build, test, and check code on every push, including license compliance checks.
- CJEU
Court of Justice of the European Union; its rulings interpret EU law for all Member States.
- Combined work
One work formed by merging separately licensed code, e.g. embedding a snippet or static linking.
- Compatibility
Whether two licenses allow their code to be combined and distributed together.
- Container image
A built binary snapshot (
.sif, OCI image) bundling many packages under many licenses.- Container recipe
The plain-text build instructions (
Dockerfile,.def); source code in its own right.- Copyleft
Licenses requiring distributed adaptations to use matching terms (GPL-3.0, EUPL-1.2).
- Copyright holder
Whoever holds the economic rights and can license the work: the author, employer, or assignee.
- Corresponding source
The full source and build scripts needed to rebuild the exact binaries you distributed.
- Derivative work
US term (17 U.S.C. § 101) for a work based on another; EU law says adaptation.
- Distribution
Giving copies to others outside your organisation; this is what triggers copyleft obligations.
- Dynamic linking
Loading a separate library at runtime; whether it creates a combined work is unsettled.
- Economic rights
Exclusive rights to copy, adapt, and distribute; often exercised by the employer (Art. 2(3)).
- EPL-2.0
Eclipse Public License, a weak copyleft license applying at the file/module level.
- EUPL-1.2
The European Commission’s copyleft license, available in 23 EU languages.
- Expression vs. ideas
Copyright protects your code, not the underlying algorithms, functionality, or interfaces.
- FSF
Free Software Foundation, the US non-profit that publishes the GPL family of licenses.
- GPL
GNU General Public License, the most widely used strong copyleft license.
GPL-3.0is the current version;GPL-2.0is still common.- Inbound licensing
The licenses on others’ code you bring into your project.
- JLA
Joinup Licensing Assistant, the Commission’s tool for comparing licenses.
- LGPL
Weak copyleft for libraries; your app may use other terms if it allows modifying and debugging the library.
- LLM
Large Language Model, the technology behind AI assistants such as ChatGPT, Copilot, and Claude.
- Memorization
When an AI model reproduces training code verbatim; that code keeps its original license.
- Mere aggregation
Separate programs shipped side by side; copyleft does not spread between them.
- MIT
The most widely used permissive license: short, simple, and requires only that the copyright and license notice be kept.
- MPL-2.0
Mozilla Public License, a weak copyleft license applying per file: modified MPL files stay MPL, new files can use any license.
- OCI
Open Container Initiative, the standard format for container images used by Docker, Podman, and registries.
- OpenRAIL
Behavioral licenses for AI models that forbid specific harmful uses; because they restrict use, they are not OSI open source.
- Originality threshold
A program is protected only if it is the author’s own intellectual creation (Art. 1(3)).
- OSI
Open Source Initiative, the non-profit that approves licenses as meeting the Open Source Definition.
- Outbound licensing
The license you choose for your own project.
- Permissive
Licenses allowing any reuse if notices are kept (MIT, Apache-2.0, BSD).
- Reciprocity
The copyleft requirement to share adaptations under matching terms.
- REUSE
FSFE standard for per-file license declarations;
reuse lintchecks they exist, not that they are correct.- RSE
Research Software Engineer: a professional who develops and maintains software used in research.
- SPDX identifier
Standard license tag in file headers, e.g.
MIT,GPL-3.0-or-later.- SPDX version suffixes
-onlyand-or-later:GPL-2.0-onlycannot move to GPL-3.0;GPL-2.0-or-latercan.- Static linking
Copying library code into your binary at build time; generally creates a combined work.
- Sublicense
Passing on permissions under your own terms; allowed by MIT, generally not by copyleft.
- Sui generis database right
EU right protecting databases built with substantial investment; unclear for model weights.
- Viral / infectious
Misleading slang for copyleft; it does not spread by mere contact.
- Weak copyleft
Reciprocity limited to a file (MPL-2.0) or library (LGPL).