Trinetra
Open console

Research

What was tried, what was measured, what failed

This is the repository's own evidence, grouped by problem area. Every source named here was actually acquired and evaluated; every figure was recorded from a run whose conditions are stated beside it. Where a candidate was rejected, the rejection is here too — a negative result is evidence.

No author, venue, year or identifier appears on this page unless it is recorded in a document in this repository. Where an area has no real reference behind it, the page says that in those words.

Source → monitoring enclave

Areas covered
7
Public corpora used
5, with per-file hashes and licences recorded
Rejected candidates
3, each with its gate result

Seven areas, six with material and one without

Grouped by the question being asked rather than by the algorithm used, because the question is what determines whether the evidence transfers.

C2, botnet and DGA

C2 beaconing and DGA are two different questions asked of the same traffic. The repository holds material for both, and it is where most of the honest negative results live: two calibrated DGA candidates trained on public corpora were evaluated and rejected rather than promoted.

Sources

Sources used for C2, botnet and DGA
FieldValueNote
IoT-23Garcia, Parmisano & Erquiaga (2020) · DOI 10.5281/zenodo.4743746Released under CC-BY-4.0. Individual conn.log.labeled files for malware captures 8, 20, 21, 34, 42 and 44 — 44,706 original flows. Supplied the primary supervised C2 labels, with raw labels retained and both det_label and detailed-label header variants decoded.
CTU-13 NerisCTU-Malware-Capture-Botnet-43 · 36,261,479 bytes40,198 normalized events in transfer replay. Used separately for raw PCAP transfer evidence. Broad botnet traffic is explicitly not labelled wholesale as C2.
CTU Normal capturesCTU-Normal-20 · CTU-Normal-21 · CTU-Normal-22282,415,864 / 311,638,284 byte benign PCAPs and the original conn/dns/ssl Bro logs for Normal-22, giving 18,892 / 10,662 / 76,166 normalized events.
Chrmor DGA corpus25 Netlab-360 DGA families plus Alexa benign namesAn independent real-world corpus, used specifically because an algorithm-derived wordlist has no held-out generator to test generalisation against. 557,717 eligible rows. Three families — necurs, kraken, symmi — were held out entirely.
Public DGA wordlistsandrewaeva/DGA · GNU GPL v2Conficker, Cryptolocker, Matsnu, Pushdo, Ramdo, Rovnix, Tinba and Zeus. At most 25,000 distinct domains per family with family metadata retained. These are algorithm-derived wordlists, not observed infection telemetry.
Benign rankingTranco 64XQX · fixed date 2026-09-141,000,000 ranked domains, 200,000 distinct accepted domains before eligibility filtering. Cite Le Pochat et al. and the fixed list. Constituent terms include Majestic CC-BY-3.0, CrUX CC-BY-SA-4.0 and Radar CC-BY-NC-4.0; no blanket commercial permission is inferred.

Recorded results

  1. DGA candidate on the public wordlist corpus: precision 0.974264, recall 0.739983, F1 0.841114, FPR 0.0198822, PR-AUC 0.975093, ROC-AUC 0.976351 over 281,847 eligible rows. Unseen-family recall 0.15914. Rejected — recall is below the 0.80 governance floor, with no exception requested.
  2. DGA candidate on the independent real-world corpus: precision 0.974928, recall 0.625296, F1 0.761916, FPR 0.017380 over 557,717 eligible rows. Unseen-family recall 0.311694. Also rejected on recall.
  3. The measured finding that matters: on the independent corpus the two transfer probes that condemned the first candidate both pass. That isolates the first failure as generator memorisation of a synthetic-lab corpus rather than a ceiling of the 14 lexical features.
  4. C2 candidate on IoT-23 labels with CTU benign captures: capture-disjoint splits of 16,275 / 14,072 / 21,997, with only 26 final positives. Every trial collapsed to one calibrated validation score; recall 0, precision 0, PR-AUC 0.001182. Rejected — the classifier detects nothing.
  5. A GRU challenger was not trained. PyTorch is absent and 26 positive test windows with a failed baseline do not support a credible promotion, so no GRU score or comparison is claimed.

Limit CSE-CIC-IDS2018 was not downloaded. The public run does not resolve temporal mismatch between modern benign rankings and older DGA algorithms, benign-proxy contamination, narrow device coverage, or the extreme final-class imbalance in the C2 set. A rejected candidate is still a measured result — the point of recording it is that the next attempt does not repeat it.

Drawn from docs/DATASET_PROVENANCE.md · docs/ML_EVALUATION.md · C2_MODEL_CARD.md · DGA_MODEL_CARD.md

Passive analysis of encrypted traffic

Encrypted sessions are analysed from handshake metadata only. There is no decryption anywhere in the system, which means the whole problem is fingerprinting and corroboration rather than content inspection.

Sources

Sources used for Passive analysis of encrypted traffic
FieldValueNote
JA4/JA4S fingerprint listr3m0s/malicious-ja4-database · AGPL-3.0 · 29 rowsPinned revision and SHA-256 recorded in vendor/ja4-intel/SOURCE.json. The committed CSV's fixed hash is verified before the local matcher is constructed.
JA4 extraction codevendor/ja4-zeek — pinned revision and upstream archive inventoryClient extraction only, BSD-licensed. docker/ja4-client.zeek loads the client JA4 and common utilities, not the JA4+ suite. Original licences are retained.
Handshake metadata readJA4 · JA4S · SNI · TLS version · ciphers · certificate attributes · size and timing envelopesProduced by the feature engine's TLS vector. No payload content is read, stored or logged at any severity above debug.

Recorded results

  1. A rare fingerprint on its own does not fire. A match requires either repetition across at least two rows of the source list, or agreement between the client and server fingerprints.
  2. Confidence is evidence-weighted and capped: 0.6, plus 0.2 when both sides agree, plus 0.15 when the list repeats the entry, never above 0.90. Severity is high when both conditions hold, medium otherwise.
  3. Client JA4 is retained under origin-only capture; server JA4S and certificate evidence require observed responses and are ignored otherwise.
  4. A benign QUIC Initial capture produces real Zeek QUIC, SNI, ALPN and client-JA4 metadata.

Limit There is no labelled JA4/JA4S corpus in this repository, so precision, recall and false-positive rate are not measurable for this detector at all — promotion required an explicit acknowledgement of that. The QUIC result is a metadata control, not a QUIC malware classification, and QUIC malware detection is not established. A fingerprint match is an investigation lead: legitimate software and firmware share TLS libraries, and the source list itself contains repeated published entries.

Drawn from docs/DETECTORS.md · docs/RUNTIME_MODELS.md · vendor/ja4-intel/SOURCE.json

Behavioural anomaly detection

Anomaly detection is the easiest place to produce a flattering number and the hardest place to produce an honest one. Both attempts in this repository are recorded, including the one that failed badly on public data.

Sources

Sources used for Behavioural anomaly detection
FieldValueNote
Isolation Forestanomaly 20260928-1 · six host-window featuresflows_per_sec_60s, bytes_out_per_sec_60s, syn_rate_5s, unique_dst_ports_60s, unique_dst_ips_60s, dns_queries_per_sec_60s. Unsupervised, so no supervised feature-importance file exists and none is fabricated.
Benign training windows2,536 windows from CTU-Normal-20Whole captures held out for validation and test; IoT-23 attack histories were evaluation-only and never used for fitting.
BaselinesPer-entity EWMA mean and variance · 256-sample quantile reservoirKeyed by entity kind, entity id and metric. A cold baseline returns none rather than a large z-score, and baseline_warm is emitted explicitly so a detector can tell within-baseline from no-baseline-yet.

Recorded results

  1. Lab champion: recorded final-test false-positive rate 0.066964 on deterministic lab-generated traffic. Not field validation.
  2. Public candidate on 6,498 held-out windows: true negatives 4,764, false positives 1,449, false negatives 285, true positives 0. Benign false-positive rate 0.233221, and 0.368608 on Normal-22 specifically. Single-row inference 6.495 ms. Rejected by precision, recall, false-positive rate and latency, with no exceptions.
  3. The measured explanation: benign browsing volume and device distributions do not transfer across captures, and malicious device histories are not necessarily volume outliers. No test-driven threshold change was made in response.

Limit Isolation Forest scores are anomaly evidence, not attack classes. A narrow benign reference is inadequate for production, and public per-class tunnel and exfiltration sensitivity is not verified. A rejected candidate is still a measured result — the point of recording it is that the next attempt does not repeat it.

Drawn from docs/RUNTIME_MODELS.md · ANOMALY_MODEL_CARD.md · docs/DETECTORS.md

DDoS

Volumetric and protocol floods are counted rather than modelled, which is deliberate: the arithmetic is checkable and a regression test can assert it. The repository's evidence here is its own measured fixtures.

Sources

Sources used for DDoS
FieldValueNote
Trigger conditionsSYN · UDP · flow rate · packet rate · amplification ratioAll thresholds live in configuration and are env-overridable. Corroboration from source count, source entropy and baseline exceedance raises confidence but never fires the detector alone.
Problem statement toolinghping3 for SYN and UDP floods; dnscat2 and iodine for DNS tunnellingNamed as suggested lab traffic generators in SIH PS 26145. The repository's generator produces flow metadata rather than packets, so the lab exercises the metadata path end to end.

Recorded results

  1. Measured on a synthetic 250-source SYN flood: syn_rate_1s 1 970/s, unique_src_ips_5s 254, src_ip_entropy_5s 7.99 bits, severity high.
  2. A sustained flood produces one result per emit tick. The first becomes an alert; the remainder bump the dedup count and advance last_seen rather than multiplying into hundreds of alerts.
  3. The detector fires only when a window actually contains at least 20 flows, so a rate computed from a nearly-empty window cannot trigger anything.

Limit No public labelled DDoS capture was acquired, so these are fixture measurements on named hardware and not a statement about field detection. The synthetic generator emits flow metadata rather than packet volume, so a line-rate or Mbps figure would be fabricated. Under origin-only capture the amplification ratio is unavailable and the detector never reports a fabricated one.

Drawn from docs/DETECTORS.md · docs/BENCHMARKING.md · docs/PROBLEM_STATEMENT.md

OT and one-way architecture

The one-way guarantee is the product's central claim, so it is enforced in software, tested through executed lifecycle evidence, and its trust assumptions are written down rather than left for a reviewer to find.

Sources

Sources used for OT and one-way architecture
FieldValueNote
Boundary enforcementUnbridged veth pair · all-protocol ingress DROP installed while both ends are downThe receive end moves into the sensor namespace and the pair is raised. The host end carries no production bridge membership, address or route. Sensor transmissions hit the drop; copied ingress travels the other way.
Lifecycle evidencereports/live/lifecycle.json · reports/live/teardown.jsonStartup, paused reconciliation, controller restart, sensor restart, sensor recreation and target recreation — each stage submitting marked ARP, IPv4 TCP/UDP/ICMP, experimental raw IP, IPv6 and raw Ethernet frames.
Positive controlA reachable monitored peer records ordinary traffic and receives zero marked sensor framesHost drop counters increase by at least seven at each stage. Ordinary TCP/UDP/ICMP socket attempts are retained alongside the raw frames. A real TLS alert after recreation proves receive capture and Redis publication recover.
TeardownA stopped disposable project leaves no receive pair and no monitored bridgeThe original lab stays healthy throughout. Run the proof with LIVE_ACCEPTANCE; scripts/setup_demo_mirror.sh --check inspects the running default demo.

Recorded results

  1. The sensor holds NET_RAW for Zeek capture, drops every other capability, and cannot administer the interface it reads from. Compose never attaches it to the production network.
  2. The API reads boundary state from a read-only shared volume, and a verified state requires a fresh successful controller status under five seconds old. Pausing the reconciler therefore reports unverified rather than verified.
  3. The absence of a return path is asserted by a test that parses the AST of every module touching captured data and fails if one imports a network or process module or calls eval, exec, system or popen.

Limit This is a software enforcement demonstration on Linux, not physical diode certification. The privileged host controller and the Docker administrator are trusted by construction, and a read-only socket mount does not restrict Docker API privileges. Status is an infrastructure assertion supported by peer-capture proof, not a cryptographic attestation. A TAP, SPAN or real diode deployment needs its own boundary and lifecycle verification.

Drawn from docs/NETWORK_ISOLATION.md · docs/SECURITY_MODEL.md · docs/LIVE_ACCEPTANCE.md

Telemetry standards

What TRINETRA can see is bounded by what the capture method exposes. Those bounds are declared in the schema rather than discovered at analysis time, which is what lets a detector know the difference between an absent measurement and a zero.

Sources

Sources used for Telemetry standards
FieldValueNote
Metadata sourcesZeek conn, dns and ssl logs correlated by uidPlus TLS client fingerprint attributes from a pinned Zeek script. Historical captures take the same path through replay, with no network involved at all.
Flow export protocolsNetFlow · IPFIX · sFlowRepresented in the event schema's source_type enumeration. This is a declared capability of the contract, not a claim that a live collector is provisioned for each protocol.
Versioned contractsEvent 1.1.0 · Feature 1.2.0 · Alert 1.1.0Independently versioned. An unknown schema version is a validation error, not a best-effort parse, and bumping the feature version invalidates every trained model — the registry refuses one recorded against a different version.
MetricsPrometheus exposition, label names fixed at constructionAn address as a label value is an unbounded time series and a way to take down the scrape target; passing an undeclared label raises. A test asserts no reserved documentation address appears anywhere in the output.

Recorded results

  1. Origin-only capture drops response-derived metadata, server fingerprints and unavailable certificate evidence. Absent observations are not zero.
  2. The response fields on a normalized event are optional precisely because on a truly unidirectional tap the reverse direction may be invisible, and detectors are written not to assume they are present.
  3. Free-text fields are length-capped and list fields count-capped before they ever reach a detector, and rejection reasons are bucketed to a small fixed set before reaching metrics.

Limit Internal data-diode deployments typically see one direction only, so a substantial part of the feature set is unavailable by construction and a detector's behaviour under that absence is a design constraint rather than an edge case. Source-type enumeration support is not the same as a live collector for every named flow-export protocol, and the repository says so rather than implying coverage.

Drawn from docs/ALERT_SCHEMA.md · docs/SECURITY_MODEL.md · docs/ARCHITECTURE.md

MITRE ATT&CK mappings

Requested area. This one has no material behind it, so it is stated rather than filled.

No reference is recorded in this repository.

No MITRE ATT&CK technique identifier, tactic or mapping appears anywhere in this repository. The threat classes the detectors emit are TRINETRA's own vocabulary — DDOS, RECON, C2_BEACON, DGA, DNS_TUNNEL, TLS_MALWARE, EXFILTRATION — not ATT&CK technique IDs, and no layer translates between the two. Publishing an ATT&CK mapping here would mean inventing identifiers, and an invented mapping is worse than an absent one: it looks like interoperability that has not been built. If a mapping is added later it belongs in the detection code, where it can be tested, not in a page.

Verified against the repository on 2026-09-30 — docs, detector source, schemas and model cards contain no ATT&CK references.

Leakage control

A precision figure is only worth what the split policy behind it is worth. These are the policies, stated, so a reader can judge them rather than take the number on trust.

Dataset split policies and the feature-change contract
FieldValueNote
Family-aware lab splitStratified 70/15/15 within each training family, plus one family excluded from training entirely and scored separately as unseen-family.A family-aware synthetic split is not evidence of host-disjoint field validation.
Temporal feedback split70/15/15 by alert timestamp.A model that only works on data from before its own training window is not useful.
Shared feature generationThe same code path produces training rows and inference rows.Changing a feature requires bumping the feature schema version, and the registry then refuses to load any model trained against the old one rather than silently mis-feeding it.
Final test is lockedThe final test split was locked before it was read and was not used for retraining.Applied to the DGA public-corpus runs, including the rejected candidates.

Sources that could not be obtained

Named here rather than left out. Each of these is a place where the evidence base is thinner than it would ideally be, and knowing which is which is the point of listing them.

Referenced sources that were not obtained, and why
FieldValueNote
DGArchiveCredentials and institutional vetting required. No access was requested.
Netlab DGA feedThe certificate had expired. Verification was not bypassed.
CTU-Normal-22 raw PCAPExceeded the per-file and disk budget. The official raw Bro logs were used instead.
Full IoT-23 archive21.5 GB exceeds available disk. Individual labelled log files were used.
CSE-CIC-IDS2018Not downloaded. Available CTU benign captures were used instead.
Public DNS-tunnel and exfiltration capturesNot acquired. Those two detectors are verified against lab fixtures only.

Absence of evidence is not evidence of absence

From evidence to behaviour

Two places to go from here.

Capabilities

What the nine detectors actually look for, with their real trigger conditions and the limit each one carries.

Read the detector specifications →

Architecture

Where every one of these measurements happens — twelve layers, a one-way boundary, and a governance lane with a person in it.

Follow the data path →