Skip to main content

Architecture Design

FieldValue
Document IDADD-OCSA-001
Version2.0
StatusBaseline — includes static and dynamic behavioural views
Date2026-09-12

Change log v1.0 → v2.0: added section 6 (static structure: module dependency graph and data-model class diagram) and rewrote section 7 (dynamic behaviour) as 13 modelled behaviours — 4 sequence diagrams, 6 flowcharts, 3 state machines — each with textual analysis.

1. Introduction​

Purpose​

This document describes the internal architecture of OCSA: module decomposition, static structure, dynamic behaviour, deployment topology and the architectural decisions that shaped them. It is the reference for anyone maintaining or extending the system.

Scope​

Phase 1 as delivered. Phase 2 extensions are noted where they affect current structure.

2. Architectural goals and constraints​

IDGoalHow achieved
AG-1Deterministic, reproducible verdictsEngine is pure functions; no I/O, no randomness in the verdict path
AG-2One verdict across all surfacesUI, CLI and MCP call the same engine functions
AG-3Zero installation for the reviewerStatic site, no build step for development
AG-4Supplier data never leaves the machineNo network calls; all analysis client-side
AG-5Auditable outputEvery finding carries rule ID + evidence path
AG-6Deployable to a CDNdist/ is a self-contained static folder
AG-7Extensible license knowledgeDeclarative rule tables in license-db.js
AG-8Survivable without maintenanceZero runtime dependencies
IDConstraintConsequence
AC-1No server-side component in Phase 1No persistence, no multi-user, no workflow
AC-2Browser sandboxNo filesystem writes from the UI; exports use Blob download
AC-3Corporate workstationNode 18+, no admin rights
AC-4Engine must be I/O-freeFile reading happens in the adapters (CLI/MCP/UI fetch)

3. System context​

External actors: supplier (input), reviewer (operator), review board / OEM (consumer), legal (authority), CI pipeline (automated gate), AI agent (MCP consumer).

4. Container view​

ContainerTechnologyResponsibilitySize
Web dashboardHTML + CSS + ES modules in browserInteractive analysis and visualisationindex.html 141 LOC, styles.css 160 LOC, app.js 397 LOC, graph.js 244 LOC
CLINode.jsBatch analysis, report generation, CI gatecli.mjs 144 LOC
MCP serverNode.js, JSON-RPC over stdioAgent/LLM interfacemcp-server.mjs 321 LOC
Domain engineES modules, pureAll analysis logicspdx.js 208, license-db.js 397, risk-engine.js 371, diff.js 97, report.js 268 = 1 341 LOC
Test suiteNode (+ jsdom, dev-only)138 assertionssmoke-test.mjs 90, ui-smoke-test.mjs 117
Static buildNodeProduce deployable dist/build-static.mjs 38 LOC

Total: 1 982 LOC in src/, 710 LOC in tools/ + scripts/, 301 LOC of markup and CSS.

5. Component view​

5.1 src/spdx.js — ingestion​

AspectDetail
ResponsibilityParse SPDX 2.x JSON into a normalised model; build the dependency graph; assess SBOM completeness
Depends onnothing
Key exportsparseSpdx(json), buildDependencyGraph(sbom), checkSbomQuality(sbom)

parseSpdx normalises each package to a flat record. Non-package elements (SPDXRef-Document-, SPDXRef-File-) are filtered out. Roots come from documentDescribes, falling back to nodes without an incoming dependency edge.

buildDependencyGraph converts relationships to directed edges, classifying each by linkage (see Linkage classification). Duplicate edges between the same pair collapse, preferring an explicit STATIC_LINK/DYNAMIC_LINK over a bare DEPENDS_ON.

checkSbomQuality runs the 12 checks described in SBOM Quality Checks and returns { checks, passed, total, score }.

5.2 src/license-db.js — license knowledge​

AspectDetail
ResponsibilityLicense classification, obligation catalogue, expression parsing, compatibility
Depends onnothing
Key exportsTIERS, OBLIGATIONS, LICENSE_RULES, classifyLicense, applyException, inspectLicenseText, parseExpression, evaluateExpression, checkCompatibility

Rule tables are declarative, matched longest-prefix and case-insensitively. Tier ranks are deliberately non-contiguous so a rank lookup is unambiguous — see License Tiers.

The classification cache maps lowercased id → classification including the computed rank. The first implementation cached the object before computing the rank, so the second and later lookups of a license returned NaN → UNKNOWN. That was BUG-03, the only critical defect, and it was silent and order-dependent.

5.3 src/risk-engine.js — analysis​

AspectDetail
ResponsibilityEffective-license resolution, copyleft propagation, conflicts, findings, scoring
Depends onlicense-db.js, spdx.js
Key exportanalyze(sbom)

Pipeline: resolve → evaluate → propagate → conflicts → aggregate → score.

CopyleftLinkageSeverity
strongstatic / unknownCRITICAL
strongdynamicHIGH
weakstatic / unknownMEDIUM (→ LIC-009)

Every ancestor is marked tainted; findings are emitted only when the target is a distributed-unit root. See Risk Scoring for the formula and Finding Rules for the catalogue.

Finding structure — { id, rule, severity, title, component, componentId, detail, obligation, action }.

5.4 src/diff.js — milestone comparison​

AspectDetail
ResponsibilityCompare two analysed SBOMs
Depends onlicense-db.js (for TIERS)
Key exportsidentityKey, diffSboms, diffVerdict

identityKey(c) is the purl with the version stripped, else name:<lowercased name>. Findings are matched on rule | component | title so regenerated SPDXIDs between releases do not create phantom churn.

5.5 src/report.js — artefact generation​

All renderers are pure string builders — no I/O, so the same functions serve CLI file output and MCP tool results. Exports: groupByObligation, renderComplianceReport, renderNoticeFile, renderSupplierInquiry, renderComponentsCsv, renderDiffReport.

5.6 src/graph.js — visualisation​

Hand-written force layout on SVG: O(n²) repulsion + spring attraction + weak centring, integrating on requestAnimationFrame. The loop stops when motion falls below a threshold (BUG-12) and restarts on interaction. Nodes above 600 are down-sampled by rank·10 + degree + root bonus + taint bonus, with the hidden count surfaced.

5.7 src/app.js — presentation wiring​

No exports; runs on load. Owns UI state ({ sbom, result, graph, selected, previous, previousName }), renders all views, handles filtering, the detail panel, and JSON/CSV export.

5.8 Adapters​

AdapterNotes
tools/cli.mjsSix commands. Boolean flags use a dedicated has() helper (BUG-07)
tools/mcp-server.mjsHand-written JSON-RPC 2.0, newline-delimited stdio. Implements initialize, tools/list, tools/call, ping. Results are cached per SBOM path
scripts/build-static.mjsCopies browser-facing files to dist/

6. Static structure​

Notation rationale​

ConcernNotationWhy
Module decompositionDependency graph (flowchart)Shows allowed import direction and layering violations
Data modelClass diagramShows entities, attributes and cardinality
Request/response over timeSequence diagramShows who calls whom, in what order
Algorithms with branchesFlowchartShows decision points and termination
Entities with discrete modesState machineShows legal transitions and guards

Module dependency graph​

Analysis. The graph is a strict DAG with three tiers and no cycles. Two properties matter:

  1. license-db.js is a sink. It imports nothing. It is the only module holding domain policy as data, so a license change touches exactly one file.
  2. All I/O stops at the adapter boundary. app.js performs fetch and FileReader; cli.mjs and mcp-server.mjs perform fs and stdio. Nothing in the engine reads a file, writes a file, or touches the network.

The second property is what guarantees AG-2: the browser, the CLI and the MCP server cannot disagree, because they execute literally the same functions on the same input. It also makes the engine testable without fixtures beyond a JSON object — which is why the engine suite needs no jsdom.

graph.js is deliberately isolated from the engine: it receives plain {id, label, color, …} nodes and knows nothing about licenses. The visualisation could be replaced without touching any analysis code.

Data model​

Analysis. Three observations:

  1. Component extends Package. Analysis does not build a parallel object graph; it decorates the parsed package with verdict fields. The original SPDX fields stay available, so a reviewer can always see why a license was chosen.
  2. Finding is immutable and self-contained. Every field needed to act is in the finding itself. Findings are the unit of export, of diffing, and of waiver (Phase 2).
  3. Verdict retains every term, not just the resulting tier. That is what lets the UI show "elected LGPL-2.1 out of LGPL-2.1 OR GPL-2.0" — the audit trail for ADR-007.

7. Dynamic behaviour​

Behaviour inventory​

IDBehaviourNotationWhere it is documented
B-01Web SBOM analysisSequencebelow
B-02CLI release gateSequencebelow
B-03MCP tool callSequenceMCP Server
B-04Milestone diffSequencebelow
B-05Effective license resolutionFlowchartDomain Concepts
B-06SPDX expression evaluationFlowchartLicense Tiers
B-07Copyleft propagationFlowchartFinding Rules
B-08Compatibility conflict detectionFlowchartbelow
B-09Risk scoringFlowchartRisk Scoring
B-10Release gate decisionFlowchartbelow
B-11Component classification lifecycleState machinebelow
B-12UI application lifecycleState machinebelow
B-13Graph simulation lifecycleState machinebelow

B-01 — Web SBOM analysis​

Analysis. Four characteristics are worth calling out.

Everything is synchronous after the file read. There is no server round-trip, so there is no loading state, no retry logic, and no failure mode beyond a parse error. For a 42-component SBOM the whole pipeline completes in well under a second; the only asynchronous part is the graph's animation loop, which is purely visual and cannot produce a wrong verdict.

Parse errors terminate early and loudly. A non-SPDX JSON is rejected at parseSpdx with a targeted message rather than producing an empty dashboard that looks like a clean result. A silently empty result would be the worst possible failure for a compliance tool.

The engine is called exactly once. analyze returns everything the UI needs. Re-rendering a filtered view never re-runs analysis, so filtering is instant and cannot change a verdict.

The graph is constructed last and cannot block the verdict. If createGraph were to fail, the findings and inventory are already on screen. Visualisation is never on the critical path to an answer.

B-02 — CLI release gate​

Analysis. The contract is deliberately narrow: the exit code is the interface. A CI system needs nothing else. This is why the gate is a separate command rather than a flag on analyze — a separate command can own a clean exit-code contract without ambiguity.

  • Reasons go to stderr, the success line to stdout. A pipeline that captures only stdout still fails on a non-zero exit, and a human reading the log sees the reasons without them being swallowed by a success message.
  • Unresolved licenses always fail unless explicitly waived with --allow-unresolved.

B-04 — Milestone diff​

Analysis. The diff operates on analysis results, not on raw SBOMs. It compares verdicts, so a change is reported when the compliance consequence changes, not merely when a string differs. If a supplier reformats an SPDX expression without altering its meaning, no diff is produced.

The identity key is load-bearing: matching on purl-without-version means qt@6.5.2 → qt@6.5.3 is one version change, whereas matching on name+version would produce one removal and one addition — and would lose the license-change comparison entirely.

B-08 — Compatibility conflict detection​

Analysis. The candidate set is filtered before comparison. NOASSERTION and LicenseRef-* are dropped because there is nothing to compare against — flagging them would duplicate LIC-003 and LIC-004 rather than add information.

The rule table is coarse: it matches on license family prefixes rather than exact identifiers, because the well-known incompatibilities are family-level properties. Each rule carries a rationale surfaced verbatim in the finding.

The known weakness is granularity: the closure is computed over the whole subtree of a root, so two components in separate executables that share a root would be reported as conflicting. This is RR-03 and P2-06 — resolving it needs a process/executable boundary map from the supplier.

B-10 — Release gate decision​

Analysis. All four checks are evaluated before any decision is taken — the gate collects all reasons rather than short-circuiting on the first. An engineer reading a CI log gets the complete list of what must be fixed, not one item per run.

Defaults are --max-critical 0, --max-high 0, --min-quality 0, the strictest sensible posture. A programme that wants a looser posture must state it explicitly in the pipeline command, and that statement is visible in version control — which is exactly where a compliance policy should be recorded.

B-11 — Component classification lifecycle​

Analysis. The states are not stored as a field — there is no state property — but the progression is real and strictly ordered.

Ordering guarantees a component is never partially classified. pickLicense always produces an effective license (falling back to NOASSERTION), so evaluateExpression always has input and tier is always assigned. There is no "unclassified" resting state, which removes an entire class of null-handling bugs.

The Tainted → FindingRaised / TaintOnly split is the fix for BUG-06. Both transitions record the taint; only the root transition produces a finding. An intermediate library is not itself an action item, but a reviewer drilling into it should see that it carries copyleft.

B-12 — UI application lifecycle​

Analysis. The state is genuinely simple — one analysis result at a time, plus an optional baseline — and the diagram deliberately shows that. There is no router, no history stack and no server session, which is why the whole UI fits in one 397-line module without a framework.

  • Loaded → Loaded on filter re-renders views but never re-runs analyze. The verdict is computed once and is immutable for the session.
  • Comparing is orthogonal to Detail. A baseline only adds a derived view (the diff card); it does not change the current analysis.

B-13 — Graph simulation lifecycle​

Analysis. This state machine exists because of BUG-12. The first implementation started a requestAnimationFrame loop and never stopped it, so an idle dashboard consumed CPU indefinitely — unacceptable for a tool left open on a workstation all day.

Settled is the important state. The loop cancels itself when the layout stops changing. Any user interaction calls wake(), which raises the simulation temperature and restarts the loop. The visual effect is identical to a permanently running simulation; the cost is not.

Stopped is reached on every re-render. When a filter changes, app.js calls stop() before creating a new graph, so discarded simulations cannot keep animating against a detached DOM. Forgetting this would leak a rAF loop per filter change.

8. Deployment architecture​

EnvironmentFormNotes
Workstationpython -m http.server (or any static server)Must be HTTP; ES modules are blocked on file://
CInode tools/cli.mjs gate …Non-zero exit blocks the pipeline
Corporate webCloudflare Pages serving dist/Custom domain + Zero Trust Access; also protect *.pages.dev
Agentnode tools/mcp-server.mjsLaunched by the MCP client; local process

_headers sets X-Content-Type-Options, Referrer-Policy, X-Frame-Options, and Cache-Control: no-cache for HTML/JS — because with no content hashing, aggressive caching would serve stale code after a redeploy.

9. Cross-cutting concerns​

ConcernApproach
ConfidentialityNo network calls anywhere; verified by static scan
DeterminismNo Date.now(), no randomness, no I/O in the verdict path
Error handlingParse failures surface a targeted message; UI shows an alert, CLI exits 2, MCP returns isError
PerformanceO(n²) layout acceptable to ~1 000 nodes; 600-node render cap; graph loop settles
Security (deployed)No user input reaches an interpreter; innerHTML output is escaped via esc(); static hosting has no attack surface
MaintainabilityRule tables are declarative; adding a license is one line
InternationalisationEnglish only; license identifiers are language-neutral

10. Architecture decision records​

See Decision Records for the full ADR table with status.

11. Extension points​

ExtensionWhereEffort
New license / exception / conflictLICENSE_RULES, LINKING_EXCEPTIONS, COMPAT in license-db.jsMinutes
New finding rulerisk-engine.js, next free LIC-0xx (LIC-002 is reserved)Hours
CycloneDX inputNew parser producing the same normalised model as parseSpdxDays
Enrichment (OSV, ClearlyDefined, deps.dev)Post-parse step keyed on purl; must remain optional and offline-capableDays
New export formatNew renderer in report.jsHours
Hosted APIWorker with a fetch handler wrapping risk-engine.js — the stdio MCP server cannot run on Workers (no stdin, no fs)Days
Persistence and workflowBackend service; engine unchangedWeeks

12. Technical debt and known limitations​

IDItemImpactPlanned
TD-01No persistence — every session re-analyses from scratchCannot diff without keeping files manuallyPhase 2 (P2-01)
TD-02No waiver/approval workflowWaivers tracked outside the toolPhase 2 (P2-02)
TD-03Compatibility conflicts lack a process-boundary mapOver-reports conflictsPhase 2 (P2-06)
TD-04O(n²) graph layoutSlow beyond ~1 000 nodesAcceptable; add Barnes-Hut if needed
TD-05UI state is module-global in app.jsHarder to unit-test view logicAcceptable at current size
TD-06Tier table not legally ratifiedVerdicts are engineering triagePhase 2 (M-8)
TD-07app.js mixes rendering and stateSome duplication between tablesRefactor if UI grows
TD-08No i18nEnglish onlyIf non-English users join
TD-09Expression parser does not validate against the SPDX license listA typo'd identifier classifies as UNKNOWN rather than erroringAdd list validation
TD-10Scoring weights are judgement, not calibrationScore is a trend indicator onlyFit against historical data if available