Build It Yourself
The mistake most people make is starting with the graph. The graph is the easy part. The hard part is encoding your compliance rules so they are deterministic and auditable. Build in this order.
Phase 0 — Freeze the rules before writing code (1–2 days, no code)
Do this first, in a spreadsheet, before you open an AI coding tool. If you skip it, the AI will invent a risk model and you will not be able to defend the result to an OEM auditor.
- License tier table. For each license family define: tier, copyleft strength, and the
concrete obligations it triggers. Start from
src/license-db.js— it already encodes strong / network / weak / file-level copyleft plus the linking exceptions that matter in embedded (u-boot-exception-2.0,GCC-exception-3.1,Classpath-exception-2.0,Linux-syscall-note). - Your release gate. What blocks a milestone sign-off? A defensible default: no unresolved license, no strong copyleft in a statically linked proprietary binary, no AGPL anywhere in a connected ECU, no known-incompatible license pair in one executable.
- The output you must produce. Usually: per-component obligation checklist, NOTICE file content, source-offer list, and a signed exception register for anything waived.
Deliverable: a Markdown/CSV rules file. This becomes the spec you paste into the AI tool.
Phase 1 — Ingest and normalise (half a day)
Read an SPDX 2.3 JSON file. Build a normalised component model with: SPDXID, name, version, supplier, downloadLocation, copyrightText, purl (from externalRefs), checksums, licenseConcluded, licenseDeclared, licenseInfoFromFiles. Build a dependency edge list from
relationships, keeping the relationship type, and classifying each edge as static / dynamic / unknown / dev / optional. Find the roots fromdocumentDescribes, falling back to nodes with no incoming edge. Also parsehasExtractedLicensingInfos. Ignore packages whose SPDXID matchesSPDXRef-(Document|File)-. Also implement the NTIA minimum-elements check and report which elements are missing.
Done here: src/spdx.js.
Non-negotiable supplier requirements — put these into your supplier quality agreement, because without them no tool can give you a defensible answer. See Supplier SBOM Requirements.
Phase 2 — The risk engine (1–2 days)
This is where all the domain value lives. Order matters:
- Expression parser. SPDX expressions are
A AND (B OR C) WITH D. Write a small recursive-descent parser. Rule:AND= worst case (all apply),OR= you may elect the most favourable — but flag the election so it gets documented. - Classification with exceptions.
GPL-2.0-or-later WITH u-boot-exception-2.0is not the same risk asGPL-2.0-only. Handle this or you will flood the report with false positives and lose credibility with the supplier. - Propagation. For every copyleft component, walk up the dependency graph. Static or undeclared linkage plus strong copyleft means the parent is a derivative work. Weak copyleft (LGPL) means the parent is not relicensed, but you owe a relinking mechanism. Mark every ancestor as tainted so the graph can show it.
- Compatibility conflicts. GPL-2.0-only + Apache-2.0, EPL + GPL, CDDL + GPL cannot be combined in one work. Report per distributed unit.
- Findings, not scores. Each finding = rule id + severity + evidence (the dependency path) + the obligation + the recommended action.
Implement a copyleft risk engine over the model from step 1. For each component resolve the effective license as licenseConcluded → licenseDeclared → licenseInfoFromFiles → NOASSERTION. Classify it into CRITICAL / HIGH / MEDIUM / LOW / UNKNOWN. Then propagate: for each strong copyleft component, walk up the dependency graph and mark every ancestor as tainted, using the linkage type of the edge (static = derivative work, dynamic = separate work). Emit findings with: rule id, severity, title, evidence, the concrete obligation, and a recommended action. Dedupe findings per distributed unit, not per hop.
Done here: src/risk-engine.js and src/license-db.js.
Phase 3 — Visualise (1 day)
Only now build the UI. Three views are enough; do not build more.
- KPI strip — risk score, critical/high counts, unresolved licenses, SBOM quality %.
- Dependency graph — force-directed, node colour = risk tier, red edge = static link, ring around a node = tainted transitively, click = detail panel. This is the view that makes the supplier conversation short.
- Findings table + component inventory — filterable, exportable.
Build a zero-dependency force-directed graph on SVG. Nodes carry id, label, colour, radius (by degree) and a tainted flag. Edges carry a linkage type. Support pan, wheel zoom, node drag, hover highlight of the neighbourhood, and click-to-select. Cap rendering at 600 nodes by keeping roots, high-risk and high-degree nodes, and report how many were hidden. Stop the animation loop once the layout settles.
Done here: src/graph.js, src/app.js, index.html.
Phase 4 — Productise (1–2 weeks)
The browser prototype has no persistence. Add, in this order:
- Backend (Node/TS or Python FastAPI) with the same engine imported — the engine must stay pure and headless so the CLI, the API and the UI all share one verdict.
- Storage — one row per (supplier, ECU, milestone, SBOM). Keep every SBOM: the value is in the diff between milestones.
- SBOM diff — done here (
src/diff.js,cli.mjs diff, UI diff card). - Approval workflow — per-finding state: open / waived / supplier-answered / accepted, with an approver and a reason. Auditors ask for the waiver register, not the scan.
- CI gate — run the engine in the build pipeline; fail the build on new CRITICAL findings.
- Report generation — obligation checklist, NOTICE file, source-offer package list, per-ECU compliance report.
Phase 5 — The agent layer (1 week, and only after Phase 4)
Once the engine is a clean set of pure functions, wrapping it as an agent is mechanical. Expose the engine as tools and let an LLM orchestrate:
| Tool | What it does |
|---|---|
parse_sbom(file) | parse + NTIA quality report |
assess_copyleft(sbom_id) | findings, score, obligations |
explain_finding(id) | why it fired, with the dependency path |
find_components(license=, tier=, supplier=) | inventory queries |
diff_sboms(a, b) | milestone comparison |
check_license_compatibility(list) | pairwise check |
generate_obligation_report(ecu, milestone) | NOTICE + source-offer list |
draft_supplier_inquiry(finding_id) | write the email to the supplier |
Build it as an MCP server so the same tools work in any MCP-capable client, and keep the LLM out of the verdict path — the LLM should explain and draft, never decide the risk tier.
What was learned
- Rules before code, always. If the model invents the risk model, you cannot defend it.
- Findings beat scores. A score is unauditable; an evidence trail is everything.
- Deduplicate aggressively. The first propagation implementation emitted 22 identical findings for one component. Unusable output is worse than missing output.
- Known-answer repetition catches the worst bugs. BUG-03 — a cache that returned a classification without its computed rank — was silent and order-dependent. It surfaced only because the suite evaluates the same license twice with different modifiers.
- Assert "nothing left hidden" for anything that reveals elements progressively. It caught three defects in the walkthrough that looked fine in code review.