Software licensing focusing on open source

Objectives

  • Principles of open source licensing

  • Difference between permissive and copyleft licenses

  • Frameworks for AI-generated and AI-assisted code

  • Determine the software license for your project following EU copyright framework

  • Navigate the Joinup Licensing Assistant to select a compliant license

  • Understand the licensing distinction between container recipes and container images

Limitations and context of this lesson

This lesson is designed as practical educational material for researchers and research software engineers, not formal legal advice.

  • EU directives set only minimum requirements in some areas: Member States implement them differently and may add national rules not covered here. For example, some Member States let university researchers retain ownership of the programs they write instead of applying the employer rule in Art. 2(3).

  • Institutional Context: Employment contracts, grant agreements, and university policies heavily influence software ownership and licensing choices.

  • This lesson covers only the general principles of open-source reuse, copyright scope, and software adaptation.

If you need formal guidance, the references below can help and so can legal experts, especially if your host institute has a legal services office:

Introduction: What is a Software License?

Under copyright law worldwide, software without an explicit license defaults to All Rights Reserved: nobody else may run, copy, modify, distribute, or build on your code. A software license is how the copyright holder exercises their exclusive rights, granting others permission to reproduce, distribute, modify, and sometimes sublicense the work.

Note that author and copyright holder may differ: under Art. 2(3), an employer exercises the economic rights in code written by an employee on the job, unless a contract says otherwise. The employee is still the author; the employer is who licenses it. This matters in practice, because the person choosing the license for a research project is often not the person who wrote the code.

This lesson focuses on open-source licenses. If your employment terms and institutional policy allow you to open-source the code you write, we recommend doing so. It makes you a better citizen of the research community, since others can reuse, verify, and build on your work. It also protects your future self: code your employer owns and never licenses stays locked behind All Rights Reserved when you change jobs, whereas an open license grants everyone the right to reuse it, including you.

Open-source licenses fall into three families, which differ in what they let downstream users do:

%%{init: {'themeVariables': { 'edgeLabelBackground': '#faf5ff', 'fontSize': '16px' }}}%% flowchart TB A["<b>Your code</b>"] -->|"no license"| B["<b>All Rights Reserved</b><br/>Nobody may run,<br/>copy or modify it"] A -->|"attach a license"| C{"What do you want<br/>downstream users<br/>to be able to do?"} C --> D["<b>Permissive</b><br/><i>MIT, Apache-2.0</i><br/>'Reuse freely, keep credit'<br/>━━━━━━<br/>Run &amp; modify ✅<br/>Closed product ✅<br/>Must share changes ❌"] C --> W["<b>Weak Copyleft</b><br/><i>LGPL, MPL-2.0</i><br/>'Share alike, within a boundary'<br/>━━━━━━<br/>Run &amp; modify ✅<br/>Closed product ✅ <i>cond.</i><br/>Must share changes ✅ <i>file/library only</i>"] C --> E["<b>Copyleft</b><br/><i>GPL-3.0, EUPL-1.2</i><br/>'Share alike'<br/>━━━━━━<br/>Run &amp; modify ✅<br/>Closed product ❌<br/>Must share changes ✅"] classDef green fill:#e6ffe6,stroke:#2b8a3e,stroke-width:2px,color:#1b4332; classDef amber fill:#fff3bf,stroke:#f08c00,stroke-width:2px,color:#5c3c00; classDef yellow fill:#fff9db,stroke:#f59f00,stroke-width:2px,color:#5c3c00; classDef dashed_red fill:#ffe3e3,stroke:#adb5bd,stroke-width:2px,stroke-dasharray: 5 5,color:#000000; classDef white fill:#f8f9fa,stroke:#adb5bd,stroke-width:2px,color:#212529; class D green; class W amber; class E yellow; class B dashed_red; class A,C white;

Three rules of thumb as an compliment to the diagram:

  • Copyleft only applies when you share the code. Running modified GPL code on your own machine or cluster creates no obligations.

  • “Weak” copyleft still has conditions. For example, if you ship an LGPL library inside a closed product, you must still let users modify and debug that library.

  • Copyleft licenses often don’t mix. Code under two different copyleft licenses may not be combinable, so your choice today decides who can build on your work later.

You will hear copyleft called viral or infectious. The slang is misleading: copyleft doesn’t spread just because GPL code sits next to yours in a repository or container. It only applies when you build GPL code into your own, for example by copying in a snippet. And choosing copyleft is a legitimate project decision, not a sign that a license is harmful.

This lesson covers both directions: choosing terms for software you write, and complying with terms attached to code written by others. The scenarios later work through each case.

Scope of this Lesson: What Counts as Software?

Across international frameworks (17 U.S.C. § 101 and WIPO model provisions), software is broadly defined as a set of statements or instructions used directly or indirectly in a computer to bring about a certain result. Research software goes well beyond Python scripts, so this lesson covers six asset types find the ones matching your own project, since the scenarios later map onto them:

  • Source Code original algorithms, or implementations of published methods.

  • Third-Party Integrations embedded snippets and linked libraries (static or dynamic).

  • Infrastructure as Code Ansible playbooks, Terraform configs, container recipes (Dockerfile, Apptainer .def).

  • Container Images built binary snapshots (.sif files, OCI registry images).

  • AI-Assisted Code generated or refactored with human oversight.

  • AI Prompt Templates engineered system prompts meeting the threshold of human authorship.

Motivation: Debugging a License Compliance Failure

With the three license families in mind, examine what happens when they collide inside an automated CI/CD pipeline:

%%{init: {'themeVariables': { 'edgeLabelBackground': '#faf5ff' }}}%% flowchart TB subgraph box["CI/CD License Compliance Debugging Pipeline"] A["Paste a snippet copied from somewhere"] --> A2["<b>Build Trigger:</b> Push to my-code-base"] A2 --> B["Run Compliance Scanner"] B --> C{"Check Inbound vs.<br/>Outbound Terms"} C -->|"Target license: Permissive<br/>Pasted snippet: Copyleft"| D["❌ <b>BUILD FAILURE</b> · job #142<br/>Pasted copyleft snippet blocks MIT release"] D --> E{"Select Patch Option"} E -->|"A: Keep MIT, add comment '# Originally GPL'"| F["❌ <b>BUILD FAIL</b><br/>Comments do not override license terms"] E -->|"B: Re-license repo to match the snippet"| G["✅ <b>BUILD PASS</b><br/>Your license now matches the snippet"] E -->|"C: Reimplement the functionality yourself"| H["✅ <b>BUILD PASS</b><br/>Your own expression, your own license"] E -->|"D: Delete LICENSE file to silence the scanner"| I["⚠️ <b>SCANNER PASSES LEGAL TRAP</b><br/>Still infringing, and your own code reverts to All Rights Reserved"] P["<b>Permissive</b><br/>MIT, Apache-2.0, 0BSD"] WC["<b>Weak Copyleft</b><br/>LGPL, MPL-2.0, EPL-2.0"] CL["<b>Copyleft</b><br/>GPL-3.0, EUPL-1.2"] end P -.->|"What I want for my repo"| C CL -.->|"What the pasted snippet uses"| C WC -.->|"Fine only as a separate library"| C classDef pass fill:#e6ffe6,stroke:#2b8a3e,stroke-width:2px,color:#1b4332; classDef copyleft fill:#fff9db,stroke:#f59f00,stroke-width:2px,color:#5c3c00; classDef amber fill:#fff3bf,stroke:#f08c00,stroke-width:2px,color:#5c3c00; classDef fail fill:#ffe3e3,stroke:#e03131,stroke-width:2px,color:#5c0000; classDef warning fill:#fff3bf,stroke:#f08c00,stroke-width:2px,color:#5c0000; classDef neutral fill:#f8f9fa,stroke:#495057,stroke-width:2px,color:#212529; classDef box_fill fill:#ffffff,stroke:#adb5bd,stroke-width:1px; class P,G,H pass; class CL copyleft; class WC amber; class D,F fail; class I warning; class A,A2,B,C,E neutral; class box box_fill;
  • Option D : deleting the LICENSE file makes the scanner quiet without changing anything legally. You are still distributing someone else’s copyleft code without honouring its terms, and you have now stripped your own users of any permission to use your work. A green pipeline is not a compliance result.

  • Option C works only if you genuinely reimplement the functionality without copying the original expression. As the idea/expression split above establishes, the algorithm is free to reuse the specific code is not. Reading the original closely and retyping a close paraphrase is still copying.

Limitations of AI-Assisted Licensing Advice

Modern software developers and RSEs routinely rely on AI coding assistants (ChatGPT, Claude, GitHub Copilot) to generate boilerplate, refactor functions, and answer project setup questions.

However, using these tools for legal or licensing guidance introduces a subtle risk of AI legal bias. AI models are overwhelmingly trained on US-centric web data and legal forum posts, so their outputs default almost universally to US common law concepts such as Fair Use, Work Made for Hire, and Derivative Works.

Developers working under EU statutory frameworks face a different legal reality around exceptions, ownership, and code adaptation. The clearest example is the term you will hear constantly:

  • US law (17 U.S.C. § 101) formally defines “derivative work”, and AI assistants reach for it to describe almost any code modification.

  • EU law (Directive 2009/24/EC, Art. 4(1)(b)) does not use that term at all. It grants exclusive rights over “the translation, adaptation, arrangement and any other alteration of a computer program” governed collectively as an adaptation.

  • Licenses vary: EUPL-1.2 defines “Derivative Works” in its own text as a contractual term, and GPL-2.0 used the phrase too. GPL-3.0 deliberately dropped it in favour of “modify” and “a work based on the Program”, because its drafters recognised the term means different things in different jurisdictions the same problem you face when an AI assistant uses it.

So when an AI assistant tells you a snippet creates a “derivative work”, treat that as a prompt to check the actual question under EU law: is this a statutory adaptation, or a combined work across a technical boundary? The rest of this lesson gives you that EU-aligned framework.

Standardizing In-File Declarations: SPDX Identifiers

Selecting a license is only half the job. Automated scanners and CI/CD pipelines need a machine-readable way to verify compliance per file without parsing legal text.

Managed by the Linux Foundation, SPDX provides standardized short identifiers for licenses (MIT, Apache-2.0, GPL-3.0-only, EUPL-1.2). The REUSE specification, maintained by the FSFE, builds on SPDX to define how every file should declare its copyright and license. Each file starts with two tags:

# SPDX-FileCopyrightText: 2026 Author Name <author@institute.eu>
# SPDX-License-Identifier: MIT
  • SPDX-FileCopyrightText names the copyright holder and year. This may be your institution rather than you personally; check your institution’s policy.

  • SPDX-License-Identifier names the license, using the exact SPDX identifier. Watch the suffix: GPL-3.0-only and GPL-3.0-or-later behave differently, so choose deliberately.

The tags only point to a license, so the full license text must also be in your repository. REUSE places one text file per license in a LICENSES/ folder (e.g., LICENSES/MIT.txt). Running reuse lint then checks that every file carries both tags and that every license it names has its text present.

Every scenario below shows the SPDX tagging for its asset type Python scripts, container recipes, and prompt templates each have their own conventions.

License Selection Decision Matrix & Scenario Index

Our decision framework is grounded in the European Commission’s Joinup Licensing Assistant (JLA), which sorts licenses across six criteria: Can (permissions), Must (obligations), Cannot (restrictions), Compatible (interoperability), Law (jurisdiction), and Support (governance).

The scenarios below are independent. Find the row that matches what you are actually building, jump to it, and skip the rest.

If you are…

Scenario

Typical Outcome

Example Licenses

Writing everything yourself

1. Own code

🟢 Free choice

MIT, Apache-2.0, BSD-3-Clause

Implementing a published algorithm

2. Algorithm implementation

🟢 Free choice copyleft if you want reciprocity

EUPL-1.2, GPL-3.0, AGPL-3.0

Pasting in a permissive snippet

3. Embed permissive

🟢 Stay permissive, keep notices

MIT, Apache-2.0, EUPL-1.2, GPL-3.0

Pasting in a copyleft snippet

4. Embed copyleft

🟡 Strong copyleft likely required

EUPL-1.2, GPL-3.0

Importing or linking a library

5. Link a library

🟡 Depends on which copyleft see below

GPL-3.0, EUPL-1.2, or permissive if weak copyleft

Writing a Dockerfile or .def

6. Container recipe

🟢 Free choice

MIT, Apache-2.0, BSD-3-Clause

Publishing a built image

7. Built image

⚠️ Multi-license bundle

Governed by each layer’s own terms

Using Copilot, ChatGPT or Claude

8. AI-assisted code

🟢 Free choice, verify for memorization

MIT, Apache-2.0, EUPL-1.2, GPL-3.0

Shipping prompts, weights or datasets

9. AI workflows & assets

🟢 Dual-license code vs. assets

MIT + CC-BY-4.0

This lesson covers nine scenarios; a typical session works through three or four. The rest are here for reference when your project changes.

Exercise-1: How do you work with others’ software and ideas?

Licensing-1: Which scenarios typically describe your work?

The text below can be copied to the collaborative document for an online poll:

## Question: How do you work with other's code?

**Choose many**. Vote by adding an `o` character:

- 1. Writing everything yourself
  - votes:

- 2. Implementing a published algorithm 
  - votes:

- 3. Pasting in a permissive snippet
  - votes:

- 4. Pasting in a copyleft snippet
  - votes:

- 5. Importing or linking a library
  - votes:

- 6. Writing a Dockerfile or .def
  - votes:

- 7. Publishing a built image
  - votes:

- 8. Using Copilot, ChatGPT, Claude, or similar
  - votes:

- 9. Shipping prompts, weights or datasets
  - votes:

Exercise-2: How small is a snippet?

Licensing-2: Can you decide from the number of lines?

The text below can be copied to the collaborative document for an online poll:

## Question: Which of these can you safely copy *based only on its size*?

**Choose many**. Vote by adding an `o` character:

- A. A one-line expression: `return max(lo, min(x, hi))`
  - votes:

- B. Five lines of ordinary boilerplate for parsing command-line arguments
  - votes:

- C. Three unusually written lines copied verbatim from a GPL-licensed solver
  - votes:

- D. Twenty lines you wrote independently after reading an algorithm in a paper, without looking at another implementation
  - votes:

- E. Anything under 10 lines is too small to be copyrighted
  - votes:

- F. None of the above: the number of lines alone does not decide
  - votes:


### Follow-up question

For each example, what information would you want to know before reusing or publishing the code?

Module 1: Clean Slate – Authoring Original Code & Algorithms

When writing original code or implementing published algorithms, no third-party license constrains your choice but who owns the code depends on your employment contract and national rules, so check your institution’s policy first.

Scenario 1: Authoring original code and algorithms

You wrote an original algorithm from scratch (in Python, C++, Rust, etc.). Your repository contains only your original source code and dependency specifications (requirements.txt, CMakeLists.txt, Cargo.toml).

  • Licensing Goal: You want maximum adoption and zero friction for commercial or academic reuse.

  • Legal Reality: External dependencies remain separate works. Because you only list them and have not bundled third-party code inside your repository, no inbound license terms constrain your choice. This changes if you ship dependencies together with your code, for example in an executable or container image (see Scenario 5 and Scenario 7).

  • JLA Selection Strategy: To ensure downstream users must keep your copyright notice while granting them maximum flexibility to incorporate your code into both open and proprietary software, you require Incl. Copyright without imposing share-alike conditions (leaving Copyleft/Share a. unselected).

Scenario 2: Choosing reciprocity for your own implementation

You developed a custom solver implementing algorithms from academic literature. You want any downstream improvements, extensions, or modifications to remain open-source and be shared back with the scientific community.

  • Licensing Goal: You want to enforce reciprocity (share-alike), preventing third parties from incorporating your implementation into proprietary software without sharing their modifications.

  • Legal Reality: The published algorithm itself is an unprotected idea anyone may implement it independently, as Scenario 1 and the SAS ruling establish. What copyright protects is your specific implementation. Nothing about implementing a published method forces a particular license; copyleft here is your deliberate choice to bind downstream distributors to matching terms.

  • JLA Selection Strategy: To enforce reciprocal sharing, you must mandate that downstream distributors disclose their modified source code (Disclose source) and license their adaptations or combined works under matching terms (Copyleft/Share a.).

Module 2: The Dependency Minefield – Inbound Code & Linking

Embedding third-party snippets or linking against external libraries introduces boundaries that can constrain your license choice. How far those boundaries reach depends on which license the inbound code carries.

Scenario 3: Embedding permissively licensed third-party code

You are building an RSE tool and copied a helper function or utility snippet from a third-party project licensed under a permissive license (e.g., MIT or Apache-2.0) directly into one of your source files.

  • Licensing Goal: You want to maintain a permissive default for your project while properly acknowledging and legally respecting the embedded third-party code.

  • Legal Reality: Permissive licenses explicitly grant you permission to copy, modify, and embed their code into your repository. However, embedding permissive code does not make the original third-party copyright disappear: you must preserve the original copyright notice and license terms for that specific snippet.

  • JLA Selection Strategy: Because inbound permissive code gives you maximum licensing flexibility, your overall repository can remain permissively licensed. To reflect this, require that notices are kept (Incl. Copyright) without imposing reciprocal sharing constraints (leaving Copyleft/Share a. unselected).

Scenario 4: Embedding copyleft third-party code

You are building a software tool and copied a non-trivial code snippet from a third-party project licensed under a copyleft license (e.g., GPL-3.0 or EUPL-1.2) directly into one of your source files. This is the situation behind the failed build (job #142) in the Motivation section.

  • Licensing Goal: Comply with legal requirements imposed by the inbound copyleft code while ensuring your overall repository remains legally compliant.

  • Legal Reality: Copying a non-trivial copyleft snippet into your source files creates a single combined work, so copyleft licensing generally extends to your whole project. Moving the snippet into a separate file of the same program does not change this. “Non-trivial” matters: a snippet too short or purely functional to qualify as the author’s own intellectual creation (Art. 1(3)) may not carry copyright at all. There is no word count or line count that draws this line if you are unsure, assume it is protected and either comply or reimplement.

  • JLA Selection Strategy: If you keep the snippet, your repository must adopt reciprocal sharing terms, so configure JLA to require source code disclosure (Disclose source) and reciprocal licensing (Copyleft/Share a.).

Module 3: Dependency Linking & Packaging

When software incorporates external dependencies, whether by dynamic linking, static compiling, or bundling binaries into container images licensing obligations expand beyond your own written source code. This module covers how dependency boundaries, build automation scripts, and packaged container artifacts affect legal compliance under the Joinup Licensing Assistant (JLA) framework.

Scenario 5: Linking against a GPL-licensed library

You are developing a software application that imports or links against an external software library licensed under GPL-3.0 (e.g., importing a GPL Python package or linking a C/C++ static/shared library).

  • Licensing Goal: Ensure legal compliance while using copyleft libraries as core dependencies in your software project.

  • Legal Reality: Whether linking creates a combined work is genuinely unsettled, and often has to be decided case by case. The FSF’s position is that linking a GPL library statically or dynamically creates a combined work; some legal scholars and Commission EUPL guidance disagree, particularly for dynamic linking through a stable API. Most Member States have no case law on this, so no firm general rule can be stated. The guidance below follows the conservative, widely-adopted reading.

  • JLA Selection Strategy: Under the conservative reading, linking to a GPL library means the combined program you distribute must be released under matching reciprocal terms, so configure JLA to require source code disclosure (Disclose source) and reciprocal licensing (Copyleft/Share a.).

Scenario 6: Authoring container recipes and environment specifications

You are creating a Dockerfile, Apptainer .def file, Conda environment.yml, or build recipe to automate the setup of your research environment. The recipe itself contains setup instructions, shell commands, and package lists.

  • Licensing Goal: You want maximum adoption and reuse of your build automation script so other researchers can freely adapt and build upon your workflow.

  • Legal Reality: Build recipes and configuration scripts are plain-text source code, separate from the software binaries they download at build time. The build instructions you write are your expression but note that a very short recipe (a FROM line plus two RUN commands) may be too trivial to meet the Art. 1(3) originality threshold and may not attract copyright at all. The same applies to a plain list of package names in an environment.yml. Longer, non-obvious recipes clearly do attract copyright.

  • JLA Selection Strategy: To allow anyone to reuse or adapt your container recipe without restrictions, require that your copyright notice is kept (Incl. Copyright) while leaving reciprocal requirements (Copyleft/Share a.) unselected.

Scenario 7: Distributing pre-built container images

You compiled and published a pre-built container image (e.g., pushing a compiled Docker image to Docker Hub, GitHub Container Registry, or an institutional registry, or sharing an Apptainer .sif file) containing an OS layer, runtime binaries, dependencies, and your application code.

  • Licensing Goal: Safely distribute compiled container images without violating the license terms of any software layer or binary included inside the image.

  • Legal Reality: A compiled container image is a multi-license aggregate bundle, not a single combined work. Distributing pre-built binaries makes you a distributor of every package inside, so source-availability obligations apply to the copyleft components (Linux base packages, coreutils, GPL libraries). But those packages sitting in the same filesystem as your application do not make your application a derivative of them: this is mere aggregation. Your own code keeps whatever license you chose; you simply also carry distributor obligations for the copyleft software you are shipping alongside it.

  • JLA Selection Strategy: Because a container image combines multiple distinct software components, JLA is used to evaluate constituent component obligations. When distributing compiled binaries containing copyleft layers, source disclosure requirements (Disclose source) must be fulfilled for those specific layers.

Module 4: Emerging Workflows & AI

AI-assisted development tools and machine learning models introduce unique legal challenges regarding copyright ownership, training data memorization, and behavioral restrictions. This module addresses how to license projects built with AI code generation tools and how to package research software that bundles AI models, weights, and datasets alongside source code.

Scenario 8: AI-assisted code generation

You used AI tools (e.g., GitHub Copilot, ChatGPT, Claude) to write functions, unit tests, or documentation for your research software repository.

  • Licensing Goal: Apply a permissive license (MIT or Apache-2.0) to your repository with confidence, without incurring hidden copyright infringement or copyleft obligations from code the AI model reproduced from its training data.

  • Legal Reality: Unmodified AI-generated outputs lack human authorship and are generally not eligible for copyright protection under current EU and international legal standards. Most real code, however, is a mix of human and AI contribution: you prompt, select, edit, and integrate. Where the line falls between AI output and your own work is unsettled and varies between Member States; there is no percentage or line-count threshold. The more you design, choose, edit, and integrate, the stronger your claim that the result is your work. Separately, if an LLM reproduces a substantial copyrighted code snippet verbatim from its training data (memorization), that output snippet retains its original copyright and license obligations.

  • JLA Selection Strategy: To ensure maximum adoption and academic reuse for your overall codebase, require that your copyright notice is kept (Incl. Copyright) while avoiding share-alike constraints (leaving Copyleft/Share a. unselected), supported by automated compliance checks.

Scenario 9: Packaging AI workflows, datasets, and model weights

You are developing research software that includes source code alongside trained machine learning model weights (.pt, .safetensors) and benchmark datasets.

  • Licensing Goal: Apply a clear licensing structure, with different licenses for different parts of the repository, that makes both the software source code and the non-code assets (data, weights) open and reusable under appropriate legal frameworks.

  • Legal Reality: Standard software licenses (MIT, GPL) are written for source code and fit datasets and model parameters poorly. Datasets may attract the EU sui generis database right where there has been substantial investment in obtaining, verifying, or presenting their contents. Model weights are a harder case: they are the numerical values learned during training, neither code nor a database, and whether they attract any copyright protection in the EU is genuinely unsettled. Because of this uncertainty, applying an explicit license to weights is about setting clear terms for your users, not about relying on a settled legal right.

  • JLA Selection Strategy: Use JLA to select an OSI-approved open-source license for the executable code component (Incl. Copyright selected), while using Creative Commons licenses (e.g., CC-BY-4.0 or CC0-1.0) for the dataset and weight files.

Best Practices: Attaching a License to Your Repository

Once you have selected a license using the JLA, you must officially attach it to your repository so automated scanners, package registries, and downstream users can verify your terms.

1. Adding the Root LICENSE File

Always place the full text of your chosen license in a plain-text file named LICENSE or LICENSE.txt at the root of your repository.

  • Exact Legal Text: Copy the standard text directly from spdx.org/licenses or choosealicense.com.

  • Copyright Header: Ensure you fill in the copyright year and copyright holder line at the top of the license text:

    Copyright (c) 2026 [Author Name or Institution Name]
    
  • Do Not Edit Terms: Never modify the legal wording of standard licenses (e.g., removing clauses from GPL or MIT). Edited texts are no longer the license they claim to be: compliance scanners cannot classify them, package registries may flag them, and downstream users have to get their own legal review before touching your code. If a standard license does not fit, pick a different standard license.

2. Documenting License Status in README.md

Add a dedicated License section near the bottom of your repository’s README.md file, along with a machine-readable badge:

## License

This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

3. Automated Compliance with the REUSE Standard

For multi-asset research repositories containing code, data, container build recipes, and prompt templates, follow the FSFE REUSE Initiative standard:

  1. Include License Texts: Place full license files inside a LICENSES/ directory (e.g., LICENSES/MIT.txt, LICENSES/GPL-3.0-or-later.txt).

  2. Add In-File SPDX Headers: Label every source file, build script, and prompt file with SPDX tags.

  3. Verify Compliance: Run the automated REUSE linter in your CI/CD pipeline:

# Install and run REUSE compliance check
pip install reuse
reuse lint

When reuse lint passes, every asset in your codebase carries a declared, machine-readable license that downstream users can check.

Summary: Resolving the Compliance Pipeline

When developing research software, license compliance is not an afterthought to debug at the end of a project, it is a proactive design choice. By using the Joinup Licensing Assistant (JLA) framework to align your repository license with your inbound dependencies from day one, your pipeline is far less likely to fail on a license conflict late in the project.

The diagram below illustrates how selecting a compatible license upfront ensures your code passes automated compliance checks and results in a legally sound release:

%%{init: {'themeVariables': { 'edgeLabelBackground': '#faf5ff', 'fontSize': '16px' }}}%% flowchart TB subgraph local["① What you do differently now before pushing"] direction LR A["Paste a snippet<br/>copied from somewhere"] --> L["Identify its<br/>license family"] --> S["Choose a compatible<br/>license + add<br/>SPDX headers"] end subgraph ci["② The same pipeline as before"] direction LR T["<b>Build Trigger:</b><br/>Push to my-code-base"] --> B["Run Compliance<br/>Scanner"] --> C{"Check Inbound vs.<br/>Outbound Terms"} end S --> T C -->|"terms match"| P["✅ <b>BUILD PASS</b> · job #143<br/>Compliant, reusable release"] classDef pass fill:#e6ffe6,stroke:#2b8a3e,stroke-width:2px,color:#1b4332; classDef neutral fill:#f8f9fa,stroke:#495057,stroke-width:2px,color:#212529; classDef fix fill:#e7f5ff,stroke:#1c7ed6,stroke-width:2px,color:#0b3d6b; classDef box_fill fill:#ffffff,stroke:#adb5bd,stroke-width:1px; class P pass; class A,T,B,C neutral; class L,S fix; class local,ci box_fill;

Compare this with the failing pipeline at the start of the lesson: the pipeline itself is identical. Nothing about the scanner changed the only difference is two decisions made before pushing.

Scenario Mapping Across the Pipeline

  • Choosing Your Own Terms (Scenario 1, Scenario 2 & Scenario 3): When you write original code, implement a published algorithm, or embed only permissive snippets, no inbound license constrains you the choice follows your goal. Pick permissive (MIT, Apache-2.0) for maximum adoption, or copyleft (EUPL-1.2, GPL-3.0) if you want downstream improvements shared back. Either way, preserve any third-party notices attached to code you embedded.

  • Handling Inbound Copyleft (Scenario 4 & Scenario 5): Copying a non-trivial copyleft snippet (e.g., CC BY-SA code from Stack Overflow, or a GPL fragment) creates a combined work. Linking against a copyleft library may do the same, depending on the license and the linking method. In both cases, selecting a compatible copyleft license upfront (GPL-3.0 or EUPL-1.2) satisfies the reciprocal terms and lets the scanner pass and checking the exact SPDX identifier first avoids the GPL-2.0-only incompatibility trap.

  • Packaging and Build Automation (Scenario 6 & Scenario 7): Keep plain-text build recipes (Dockerfiles) permissively licensed for maximum reuse, while annotating compiled container image binaries as multi-license aggregate bundles to satisfy embedded base-layer obligations.

  • AI Assets and Dual-Licensing (Scenario 8 & Scenario 9): Run code-similarity scanners to catch LLM training memorization before releasing AI-assisted code, and apply dual-licensing to separate executable software code (MIT) from non-code datasets and model weights (CC-BY-4.0).

  • Standardized Distribution: Adding machine-readable SPDX headers across every script, Dockerfile, and prompt template lets reuse lint confirm that every asset has a declared, documented license. Note what this does and does not prove: the linter verifies that declarations exist and are well-formed, not that they are legally correct or mutually compatible. Automation makes your intent auditable it does not replace the judgment calls in the scenarios above.

Glossary