Architecture
One path in. Nothing goes back.
TRINETRA sits in a monitoring enclave that has no path into the network it watches. Copied traffic crosses into it and is turned into evidence. The absence of a return path is not a limitation handled gracefully — it is what the system is for.
Twelve layers carry an observation from source to analyst. Every claim below names the document in this repository it came from.
Source → monitoring enclave
The whole path, once
Sources cross the boundary and never come back. The governance lane beside the pipeline is not a thirteenth layer — it branches off the alert engine and runs on analyst decisions, not on live traffic.
Data direction. Every trace points from source toward the enclave.
The one-way boundary. Crossing is inward only.
A return path, drawn so it can be seen to be stopped.
The twelve layers, in order
Each layer opens onto the technical detail behind it: the fields it reads, the thresholds it applies, and the document the claim is drawn from.
Untrusted traffic metadata enters from wherever a copy is available.
TRINETRA never reaches for traffic. A copy is supplied by whatever already sits on the link: a live TAP or SPAN port, a hardware data diode, an exported flow record, a historical capture, or a seeded lab generator. Every event carries the source type it arrived as, so a replayed capture is never mistaken for a live observation. A host on the monitored network chooses the domains, the SNI values and the byte counts in this data, so the source is treated as untrusted from the first byte.
- Source types9 recorded
- Event contract1.1.0
- Payload contentNever stored
Technical detail for layer 01 — Sources
| Field | Value | Note |
|---|---|---|
| source_type | LIVE_TAP · SPAN · DIODE · PCAP_REPLAY · ZEEK_LOG · NETFLOW · IPFIX · SFLOW · SYNTHETIC | The declared value. Source-type support is a schema capability; it does not mean a live collector is provisioned for every listed protocol. |
| flow_id | Sensor-supplied, or SHA-256 of the 5-tuple | Direction-sensitive, so a request and its response never collide. |
| event_id | Deterministic from timestamp, flow key and kind | Makes replay idempotent: the same capture ingested twice produces the same ids. |
| pcap_reference | Filename only | Captures are referenced for pivoting and never copied into the analytical store. |
Drawn from docs/ARCHITECTURE.md · docs/ALERT_SCHEMA.md
The absence of a return path toward the monitored network is the product.
Copied traffic crosses a boundary that has no path back. In the live deployment a trusted host controller provisions an unbridged receive-only interface, installs an all-protocol ingress drop on the host end while both ends are down, and then raises the pair. Sensor transmissions hit that drop. The controller reconciles the boundary every second, and the sensor reads the resulting state. This is a software enforcement demonstration on Linux — it is not physical diode certification, and the repository says so.
- Return pathNone
- Sensor capabilitiesNET_RAW only
- Reconciliation1 s
Technical detail for layer 02 — One-Way Boundary
| Field | Value | Note |
|---|---|---|
| boundary_state | HOST_BLOCKED | Requires a fresh successful controller status, less than five seconds old. Pausing the reconciler therefore reports UNVERIFIED rather than verified. |
| Enforcement | Unbridged veth pair; host ingress DROP installed before either end is raised | A missing or changed drop rule makes the controller bring the host interface down and report failure. |
| Trust limit | The privileged host controller and the Docker administrator are trusted | Stated as a gap, not hidden. A read-only socket mount does not restrict Docker API privileges. |
| Proof | reports/live/lifecycle.json · reports/live/teardown.json | Marked sensor frames are submitted at each lifecycle stage; a reachable peer records zero of them while the host drop counters climb. |
Drawn from docs/NETWORK_ISOLATION.md · docs/SECURITY_MODEL.md §1
A sensor reads the copy and publishes normalized events.
The sensor layer tails Zeek conn, dns and ssl logs correlated by uid, and publishes normalized events. A historical capture takes the same path through replay: file to Zeek to events, with no network involved at all. The live sensor holds NET_RAW for capture, drops every other capability, and cannot administer the interface it reads from. Traffic generators and attack tooling live in separate lab networks, outside the passive analysis boundary.
- ReaderZeek conn / dns / ssl
- ReplayLocal file only
- Sockets to captured hostsNone permitted
Technical detail for layer 03 — Passive Ingestion
| Field | Value | Note |
|---|---|---|
| sih_ntd/sensor/live.py | Tails live Zeek logs and publishes normalized events | |
| sih_ntd/sensor/replay.py | PCAP to Zeek to events, via subprocess on a local file | The path is resolved inside the configured PCAP directory, must carry a .pcap or .pcapng suffix and must be a regular file. Zeek argv is a list and shell= is never passed. |
| sih_ntd/sensor/synthetic.py | Seeded generator: benign load plus seven attack patterns | Flow metadata only — no packets, no sockets. Same seed, same traffic. |
| Enforcement test | test_no_module_opens_a_socket_toward_captured_addresses | Parses the AST of every module that touches captured data and fails if one imports socket, requests, urllib, httpx, subprocess, http or ftplib, or calls eval, exec, system or popen. |
Drawn from docs/ARCHITECTURE.md · docs/SECURITY_MODEL.md
Captured fields are validated, aliased and frozen into one event contract.
This is the trust boundary. Field names are resolved through a static alias table rather than evaluated, every value is validated with extra set to forbid, and a record that fails is rejected and counted by low-cardinality reason — never coerced into shape. Domains are length-checked and canonicalised, addresses go through IPvAnyAddress, ports are bounded, counters must be non-negative. A deterministic event id makes replay idempotent and a bounded dedup memory keeps a replayed capture from becoming a second copy of itself.
- Extra fieldsForbidden
- RejectionCounted, never coerced
- Event schema1.1.0
Technical detail for layer 04 — Normalisation
| Field | Value | Note |
|---|---|---|
| Validated | Domains ≤253 chars, labels ≤63, LDH charset plus _ and * · ports 0–65535 · non-negative counters | Tested against oversized names, SQL-looking strings, shell-looking strings, path traversal, markup, malformed addresses, out-of-range ports and 500-element answer lists. |
| origin_only | Removes response-derived metadata, server fingerprints and unavailable certificate evidence | The resp_* fields and response_observed are optional precisely because on a truly unidirectional tap the reverse direction may be invisible. Detectors are written not to assume they are present. |
| No payload | DnsMetadata stores names and counts; TlsMetadata stores handshake attributes | No payload content, ever. Absent observations are not zero. |
Drawn from docs/SECURITY_MODEL.md §2 · docs/ALERT_SCHEMA.md
Events become a durable, replayable event fabric between producers and workers.
Validated events are published onto Redis Streams inside the enclave: raw, normalized, features, alerts, feedback and model events, each with a dead-letter stream beside it. Consumers acknowledge only after the batch is persisted, so a crash re-delivers rather than loses. A poisoned entry is claimed back with XAUTOCLAIM and, after a bounded number of deliveries, moved to the dead-letter stream and acknowledged. Streams are trimmed with approximate max-length, which is O(1); exact trimming is not.
- Streams6 + dead-letter
- AcknowledgementAfter persist
- Poison entriesBounded, then dead-lettered
Technical detail for layer 05 — Stream Layer
| Field | Value | Note |
|---|---|---|
| Streams | network.raw · network.normalized · network.features · network.alerts · network.feedback · model.events | Each has a <stream>.dlq beside it. |
| Trimming | XADD … MAXLEN ~ 100000 | Approximate trimming is O(1). Exact trimming is not. |
| Backpressure | Bounded retries, then StreamUnavailable | No infinite loop when Redis is unreachable. |
| Queue behaviour | queue_high_water 500 · shed_shadow_total 0 | A quiet stream mid-burst does not drop a trailing burst: the pipeline flushes open windows so the tail is still evaluated. |
Drawn from docs/ARCHITECTURE.md
One vector per entity carries every window, so a detector never joins across them.
The same flow is evidence of different things depending on which end you look from, so the engine keeps three keyed perspectives: the source host, the destination, and the source-to-destination pair. Each carries 1, 5, 30, 60 and 300 second windows in a single vector, together with DNS lexical and behavioural features, TLS handshake attributes, exfiltration volume and baseline comparisons. Baselines are per-entity EWMA mean and variance with a 256-sample quantile reservoir, and a cold baseline is reported as cold rather than as a large z-score.
- Windows1 · 5 · 30 · 60 · 300 s
- PerspectivesHOST · DESTINATION · PAIR
- Feature schema1.2.0
Technical detail for layer 06 — Feature Context
| Field | Value | Note |
|---|---|---|
| Entity kinds | HOST · SUBNET · DESTINATION · SERVICE · PROTOCOL · PAIR | Alert attribution follows the same rule: a flood is attributed to the target, everything else to the source. |
| Baseline features | *_zscore · *_baseline_mean · *_baseline_p95 · *_baseline_warm | baseline_warm is emitted explicitly so a detector can tell within-baseline from no-baseline-yet. A cold baseline yields z-score 0.0, never a large number. |
| Attacker-controlled context | Carried in labels, not in values | The queried domain, JA4 and SNI are kept out of the numeric vector so it stays a clean model input. |
| Bounded state | LruStateMap, default 50 000 entities | A spoofed-source flood creates one entity per source address. Least-recently-seen entities are evicted and evictions are counted. |
| require() | Raises KeyError rather than substituting 0.0 | A silently zero-filled feature turns a broken engine into no alerts. |
Drawn from docs/ARCHITECTURE.md · docs/ALERT_SCHEMA.md
Nine detectors, one contract, each isolated from the others' failures.
Every detector implements the same contract and returns either a result or nothing, because a benign window is the normal case and should not fill the store. One detector raising is caught, counted, and the remaining detectors still see the window. A detector without a model is not silently skipped: it reports UNAVAILABLE with a reason and returns nothing. The technique is recorded on every result, so an alert never overstates how much learned machinery was involved — rules, statistics and models are all first-class.
- Detectors9, all READY
- TechniquesRULE · STATISTICAL · SUPERVISED · HYBRID
- IsolationPer detector
Technical detail for layer 07 — Detection Mesh
| Field | Value | Note |
|---|---|---|
| Rule detectors | recon · tls_malware | Fan-out and failure counted exactly. A curated, hash-verified JA4/JA4S list, not a model. A model here would replace checkable arithmetic with an unexplainable number. |
| Statistical detectors | c2_beacon · exfiltration · anomaly | Periodicity and inter-arrival statistics; volume plus baseline deviation; Isolation Forest over host-window features. |
| Supervised detectors | c2_ml · dga | XGBClassifier and CalibratedClassifierCV, both lab-trained, both on feature schema 1.2.0. |
| Hybrid detectors | ddos · dns_tunnel | Counted thresholds for the trigger plus corroborating statistics that raise confidence but never fire alone. |
| Thresholds | All in ThresholdSettings, all env-overridable | No magic numbers in a detector body. |
Drawn from docs/DETECTORS.md
Detector hits are deduplicated, corroborated and grouped into incidents.
A sustained flood produces one result per emit tick. The first becomes an alert; the rest bump the dedup count and advance last_seen. Corroboration is bounded and monotonic — confidence is the primary value plus 0.08 per distinct corroborating detector, capped at 0.99 — so it can be recomputed from the stored contributions rather than taken on trust. Two or more distinct threat classes on one subject within 300 seconds become an incident. One class is an alert; two is a story. Challenger results are dropped here and can never become an alert.
- Corroboration step0.08 per detector
- Confidence cap0.99
- Incident window300 s
Technical detail for layer 08 — Fusion
| Field | Value | Note |
|---|---|---|
| confidence | min(0.99, primary + 0.08 × distinct corroborating detectors) | Recomputable from evidence.detector_contributions. |
| Shadow results | shadow=True, persisted, dropped by fusion | Challenger output is kept for comparison and never becomes an alert. |
| Severity is separate | Impact, computed independently of confidence | High confidence in a small event stays LOW severity. Classes that imply compromise carry a severity floor; two agreeing detectors escalate one step. |
Drawn from docs/DETECTORS.md · docs/ARCHITECTURE.md
Measurement and reference are recorded separately, so a reader can tell them apart.
An alert carries what was observed and what it was compared against, in two separate blocks. Observations are the measured values. Baseline comparisons are the thresholds, baselines, z-scores and percentiles they were read against. Splitting them means a reader does not need to know the feature naming convention to tell a number from a reference. Reason codes, detector contributions, the feature schema version, the detector version, the model version and the pipeline run id travel with the record, which is enough to reproduce the decision later.
- Flow references≤128 per alert
- Reason codes≤32 per result
- Decision recordAppend-only ledger
Technical detail for layer 09 — Evidence
| Field | Value | Note |
|---|---|---|
| Evidence | summary · reason_codes · observations · baseline_comparisons · detector_contributions · flow_ids · pcap_references | feature_importances is reserved for asynchronous SHAP-style explanation and is not populated today. |
| Lineage | detector_versions · model_versions · feature_schema_version · event_schema_version · pipeline_run_id | Enough to reproduce the decision. |
| alert_ledger | Format 2, hashes verdict, evidence and lineage | Verified through /api/v1/ledger/verify. It is not an externally signed checkpoint and is not presented as one. |
Drawn from docs/ALERT_SCHEMA.md · docs/ARCHITECTURE.md
The alert is a versioned record, not a notification.
Every alert is a structured instance of one versioned schema: timestamp, first and last seen, threat class, confidence and how that confidence was derived, severity and impact, status, the subject it is attributed to, the detector and model that produced it, the evidence, the lineage and any incident it belongs to. Two internally versioned contracts govern the whole system, and an unknown schema version on an event or an alert is a validation error rather than a best-effort parse. Bumping the feature schema invalidates every model trained against the old one, and the registry refuses to load it.
- Alert schema1.1.0
- Threat classes7
- StatusesNEW → TRIAGED → CONFIRMED · DISMISSED · SUPPRESSED
Technical detail for layer 10 — Alert Engine
| Field | Value | Note |
|---|---|---|
| threat_class | DDOS · RECON · C2_BEACON · DGA · DNS_TUNNEL · TLS_MALWARE · EXFILTRATION | Anomaly evidence supports these classes; it is not itself a new attack category. |
| confidence_basis | RULE_EVIDENCE · NORMALISED_ANOMALY · CALIBRATED_PROBABILITY · FUSION | Recorded so the number is reproducible from the stored supporting features. CALIBRATED_PROBABILITY describes fitted calibration, not proven field calibration. |
| correlation_id | Target for a flood, source for everything else | Attributing a flood to one of 250 spoofed sources would mis-word the alert and scatter one incident across hundreds of subjects. |
| Known scope | Alerts contain IP addresses and domain names | No payload content is ever stored, but addresses and names are personal data in some jurisdictions. Retention is a prune command, not automatic. |
Drawn from docs/ALERT_SCHEMA.md
Read paths are open; anything that changes state needs an operator token.
The analyst surface is a REST API and a WebSocket alert stream. Readiness reports the store, the queue, the sensor and each detector separately, so a healthy process and a healthy pipeline are never conflated. Metrics are exposed in Prometheus format with label names fixed at construction — an address as a label value would be an unbounded time series and a way to take down the scrape target. Mutating routes, feedback submission, replay control and every model lifecycle action run behind an operator check first.
- Read routesUnauthenticated
- Mutating routesOperator token
- Metric labelsDeclared at construction
Technical detail for layer 11 — API / WebSocket
| Field | Value | Note |
|---|---|---|
| Endpoints | GET /ready · /health · /metrics · /api/v1/alerts · /alerts/summary · /incidents · /flows · /detectors · /models · /model-health · /drift · /feedback · /benchmarks · /ledger/verify · WS /ws/alerts | |
| require_management_access | Authorization: Bearer <token> or X-API-Key | With no token configured, mutating routes are allowed only while the API is bound to loopback. Binding to 0.0.0.0 without a token disables them with HTTP 503. |
| Known gaps | No authentication on reads, no role-based authorisation, no TLS, no rate limiting | Stated as gaps in the repository rather than left for a reviewer to find. Put the API behind an authenticated reverse proxy before broad exposure. |
Drawn from docs/SECURITY_MODEL.md §2–3 · docs/ARCHITECTURE.md
The console explains. It cannot act, because the boundary gives it nothing to act through.
Everything the console shows came across the boundary read-only: alerts, incidents, flows, detector state, model governance, replays and benchmarks. The same absence of a return path that shapes the sensor shapes the interface — there is no block button, no quarantine action and no push to a device, because there is no path to carry one. The analyst's output is a decision and, when they label an alert, a training candidate.
- Operator actions on the networkNone
- Analyst outputDecisions and labels
- Full architectureOn this page, not in the console
Technical detail for layer 12 — Analyst Console
| Field | Value | Note |
|---|---|---|
| Console | Next.js 16 App Router, React 19, reading the API through same-origin routes | |
| Absent by design | No inline path and no device-control code anywhere in the repository | TRINETRA cannot block, filter, quarantine or remediate. It cannot re-contact a source, complete a handshake or push a mitigation back across the ingest path. |
| Console routes | Overview · Alerts · Incidents · Network · Detectors · Models and learning · Replay · Benchmarks · System | This architecture page exists so the console does not have to carry the full pipeline diagram. |
Drawn from TRINETRA_DESIGN_SYSTEM.md §10 · docs/SECURITY_MODEL.md §1
Model governance is a lifecycle with a person in it
A model does not retrain itself and it does not deploy itself. Live traffic can become a training candidate and nothing more. Reaching production requires analyst labelling, dataset building, training, evaluation, governance gates and an explicit human promotion — in that order.
Analyst feedbackHuman
An analyst labels an alert TRUE_POSITIVE, FALSE_POSITIVE, FALSE_NEGATIVE or UNCERTAIN, with the corrected threat class required when the label is a false negative. The response says what it did: training candidate only, production models unchanged.
Data quality and driftSignal only
Feature distribution drift by PSI or JS divergence, prediction and confidence distribution drift, class distribution divergence, and precision decay read from analyst labels. Each finding becomes a DriftEvent carrying a metric, a score, a threshold, a window and a recommendation.
Training data factoryHuman
Two builders with different documented split policies: a stratified family-aware lab split for the DGA path, and a temporal split by alert timestamp for the feedback path. Each records its dataset id, schema version, class distribution, split sizes, label provenance, content hash, row count and seed.
Offline trainingAutomated
Fixed seed, trained on a held-out split, scored for precision, recall, F1, PR-AUC, ROC-AUC, FPR, Brier and measured single-row inference latency, plus out-of-family and out-of-corpus transfer metrics. Training registers a candidate. It never promotes.
Challenger evaluationAutomated
The challenger scores the same feature vectors as the champion. Its results are persisted with the shadow flag set, and fusion drops them, so a challenger can never produce an alert. Comparison reports held-out metrics and live score-distribution drift.
Governance gatesAutomated
Twelve gates, all of which must pass or be explicitly acknowledged with a recorded reason: precision, recall, false-positive rate, measured inference latency, feature schema match, artifact integrity, a recorded dataset, regression tests, improvement on the champion, generalisation to an independent generator, independent benign false-positive rate, and a check that the result is not implausibly perfect.
Human approvalHuman
Promotion happens only through an explicit operator command, and only after the gates have run. The gate results, the actor and the reason are appended to a promotion log that is never rewritten. A model reaching production without a person deciding it is not a supported path.
Production and rollbackHuman
Promotion to champion retires the incumbent and records the previous champion, so a rollback selects it from audit history. Models load once at worker start and are never hot-reloaded: swapping a model under a running pipeline would make alerts non-reproducible mid-stream.
The automatic chain from live network traffic to automatic production replacement is never implemented. A host that can influence the monitored network can influence the training data, and therefore the production model. Nothing here writes to a champion at runtime.
How a gate can be overridden, and where that record goes
A gate may be accepted as failing, and doing so requires a written reason. The override is appended to the promotion log alongside the gate results, marked as acknowledged, and stays visible. That is the whole mechanism: “we knew and decided anyway” has to remain auditable rather than become folklore.
Statuses move CANDIDATE to CHALLENGER to CHAMPION, then RETIRED or REJECTED. Promotion retires the incumbent and records the previous champion, so a rollback selects it from audit history. Restoration still requires that the retired binary exists and is compatible — not every historical artifact ships, and a rollback to a missing binary is not something the registry pretends to do.
What a passive system cannot do
Stated plainly, because a boundary that is only described in the positive is not a boundary anyone should rely on.
No path back — not a policy, an absence
There is no inline path and no device-control code anywhere in this repository. The absence is enforced by construction and checked by a test that parses the AST of every module touching captured data, rather than by a configuration flag someone could turn off.
- Block, filter or rate-limit anything on the monitored network
- Send a packet, probe a host or complete a handshake toward a source
- Quarantine a host or push a change to a device
- Decrypt a TLS or QUIC session
- Retrain a model and deploy it without a person deciding to
What comes back the other way is a decision and, when an analyst labels an alert, a training candidate. Read what the detectors actually look for or what has been measured and what has not.