Trinetra
Open console

Capabilities

Nine detectors, and how each one decides

Rules, statistics and models are all first-class here. The technique is recorded on every result, so an alert never implies more learned machinery was involved than actually was — forcing a model onto port-scan counting would replace checkable arithmetic with an unexplainable number.

Every threshold below is a configuration value, not a constant in a detector body. Each detector carries its own limit, stated specifically rather than as a disclaimer.

Source → monitoring enclave

Detector families
9, all READY
Techniques
Rule · Statistical · Supervised · Hybrid
Entity perspectives
Host · Destination · Pair
Windows
1 · 5 · 30 · 60 · 300 s

What each detector looks for

Threat class, entity perspective and window size first, then the counted conditions that can make it fire, the fields its evidence is drawn from, and the limit it carries.

DDOS

Distributed denial of service

HYBRIDno model
Entity
DESTINATION
Window
1 s, read against 5 s and 60 s

Counted rate thresholds fire the detector; source-distribution and baseline statistics raise its confidence but never fire it on their own. A flood is attributed to the target rather than to one of the spoofed sources, because a flood spread over 250 sources is one event and attributing it to a single address would scatter the incident.

Trigger conditions for the DDOS detector
FieldValueNote
SYN_FLOOD_RATEsyn_rate_1s ≥ 500 and syn_ratio_5s ≥ 0.5
UDP_FLOOD_RATEudp_rate_1s ≥ 1000 and udp_ratio_5s ≥ 0.5
FLOW_RATE_EXCEEDEDflows_per_sec_1s ≥ 1000
PACKET_RATE_EXCEEDEDpackets_per_sec_1s ≥ 20000
AMPLIFICATION_RATIOamplification_ratio_5s ≥ 5 and UDP-dominated

Evidence fields

  • unique_src_ips_5s
  • src_ip_entropy_5s
  • amplification_ratio_5s
  • *_baseline_mean
Limit and recorded evidence — DDOS

With no warm baseline the result carries BASELINE_COLD and severity is capped below critical — without a baseline the system cannot claim traffic is abnormal for this host. Under origin-only capture the amplification ratio is unavailable and the detector never fabricates one.

Recorded detail for the DDOS detector
FieldValueNote
Entry gateflow_count_1s ≥ 20A rate computed from a nearly-empty window cannot trigger anything.
DISTRIBUTED_SOURCES≥ 50 unique sources and ≥ 4 bits of source entropyCorroboration only — raises confidence, never fires alone.
BASELINE_EXCEEDEDObserved ≥ 10× the destination's EWMA baseline
Measuredsyn_rate_1s 1 970/s · unique_src_ips_5s 254 · src_ip_entropy_5s 7.99 bitsOn a synthetic 250-source SYN flood. Severity HIGH. A measured fixture result, not a capacity claim.
Confidence0.5 + 0.4 × normalise_exceedance(observed, threshold)Plus at most 0.10 for corroboration, capped at 0.99. The normaliser is log-scaled, so a 100× flood stays distinguishable from a 4× one.

Drawn from docs/DETECTORS.md · docs/ALERT_SCHEMA.md

RECON

Reconnaissance and port scanning

RULEno model
Entity
HOST
Window
60 s, minimum 20 flows

Fan-out needs a failure signal behind it. A wide fan-out on its own describes a busy proxy or a backup agent, so the detector requires counted ports or hosts and then requires that those connections failed. A test asserts a busy client whose connections succeed is not a scan.

Trigger conditions for the RECON detector
FieldValueNote
VERTICAL_PORT_SCAN≥ 20 unique destination ports
HORIZONTAL_HOST_SCAN≥ 25 unique destination hosts
NETWORK_SWEEPBoth of the above

Evidence fields

  • unique_dst_ports_60s
  • unique_dst_ips_60s
  • failed_ratio_60s
  • syn_ratio_60s
Limit and recorded evidence — RECON

No model, deliberately. Fan-out and failure are counted exactly; a model would add opacity and remove the arithmetic a reader can check.

Recorded detail for the RECON detector
FieldValueNote
Failure gatefailed_ratio_60s ≥ 0.8 or syn_ratio_60s ≥ 0.8This is the line between a scanner and a legitimate high-fan-out client. Without it, any wide client looks like a scan.
Regressiontest_recon_ignores_a_busy_client_whose_connections_succeedThe false-positive case is a test, not a claim.
techniqueRULERecorded on every result so confidence and technique are never confused.

Drawn from docs/DETECTORS.md

C2_BEACON

C2 beaconing, statistical

STATISTICALno model
Entity
PAIR
Window
Inter-arrival series, ≥ 6 intervals

Regularity alone never fires. A periodicity score is computed from the coefficient of variation, scaled by how much evidence exists, and at least two corroborating signals are required before an alert is raised. NTP, software updaters and monitoring agents are all excellent beacons, and that gate is what stops them becoming alerts.

Trigger conditions for the C2_BEACON detector
FieldValueNote
RARE_DESTINATIONDestination is uncommon for this host
STABLE_PAYLOAD_SIZEByte-stability across intervals
STABLE_CONNECTION_DURATIONDuration-stability across intervals
EXTERNAL_DESTINATIONOutside the internal estate
SHORT_CHECKIN_FLOWSShort, repeated exchanges
RARE_SERVICE_PORTService port uncommon for the host
STRUCTURED_PERIODICITYRepeating jitter rather than random jitter

Evidence fields

  • periodicity_score
  • interarrival_cv
  • byte_stability_cv
  • duration_stability_cv
  • destination_rarity
Limit and recorded evidence — C2_BEACON

Short high-jitter series are not established as detectable, and duration corroborators are omitted when responses were not observed. The five-second interval floor was raised from one second after a fixture run produced a false positive on ordinary sub-second uploads.

Recorded detail for the C2_BEACON detector
FieldValueNote
periodicity_score(1 − min(1, CV)) × min(1, n / min_intervals)Regularity scaled by how much evidence exists, so three lucky connections cannot score 1.0.
Interval boundsMean in [5 s, 3600 s], absolute jitter ≤ 30 s
Spectral branchRequires at least 12 intervals; autocorrelation is support-scaled
Corroborators requiredAt least 2 of the 7 listed signals

Drawn from docs/DETECTORS.md

C2_BEACON

C2 beaconing, supervised model

SUPERVISEDmodel 20260928-1
Entity
PAIR
Window
Same inter-arrival series

A gradient-boosted classifier over the origin-safe pair-window contract: interarrival mean, standard deviation, coefficient of variation and count, periodicity, burstiness, byte and duration stability, mean flow bytes, mean 60-second duration, destination rarity, observed span and an internal-destination indication.

Trigger conditions for the C2_BEACON detector
FieldValueNote
Model outputProbability above the recorded decision threshold
Pair features11 origin-safe features per pair window

Evidence fields

  • periodicity_score
  • burstiness
  • destination_rarity
  • internal_destination_indication
Limit and recorded evidence — C2_BEACON

Lab-trained. Recorded test precision 1.0 and recall 0.991667 describe deterministic lab data and do not establish field performance. The statistical C2 detector also runs independently, so a C2 alert does not by itself establish which of the two performed well.

Recorded detail for the C2_BEACON detector
FieldValueNote
EstimatorXGBClassifier
Championc2_ml 20260928-1Committed artifact, feature schema 1.2.0. SHA-256 recorded in the runtime bundle.
LoadingHash verified before deserialization, class name checked afterjoblib.load executes pickle opcodes. A file whose hash does not match what this system produced raises ArtifactIntegrityError.

Drawn from docs/RUNTIME_MODELS.md · C2_MODEL_CARD.md

DGA

Domain generation algorithm

SUPERVISEDmodel 20260928-1
Entity
HOST
Window
Per DNS query

A calibrated classifier over 14 lexical features of the queried name: length, label count, longest label, subdomain depth, payload length, suffix length, character entropy, digit ratio, vowel ratio, unique-character ratio, hexadecimal-character ratio, longest consonant run, hyphen count and digit presence. Query names are attacker-controlled, so they are carried as context rather than as numeric model input.

Trigger conditions for the DGA detector
FieldValueNote
Calibrated probabilityAbove the recorded threshold on the lexical vector
InputQuery name only — no response content, ever

Evidence fields

  • char_entropy
  • digit_ratio
  • vowel_ratio
  • unique_char_ratio
  • hex_char_ratio
  • longest_consonant_run
Limit and recorded evidence — DGA

Lab-trained and honestly scoped. In-distribution lab precision and recall are 1.0, while recorded unseen-family and independent-generator recall are 0.0. Arbitrary random or real-world DGA names are not guaranteed to trigger it. The live demo deliberately uses the lab-fixture profile because that is the distribution the model was fitted to.

Recorded detail for the DGA detector
FieldValueNote
EstimatorCalibratedClassifierCV
confidence_basisCALIBRATED_PROBABILITYDescribes the fitted calibration. It is not a claim of proven real-world calibration.
Feature contract14 lexical features, feature schema 1.2.0
Public-corpus historyTwo calibrated candidates trained on Tranco 64XQX, public family wordlists and an independent real-world corpus were both rejected on recallOne scored 0.740 test recall with 0.159 unseen-family recall; the other 0.625 test recall with 0.312 unseen-family recall, against a floor of 0.80. Neither was promoted, with no exception.

Drawn from docs/DETECTORS.md · docs/RUNTIME_MODELS.md · DGA_MODEL_CARD.md

DNS_TUNNEL

DNS tunnelling

HYBRIDno model
Entity
HOST
Window
Per query, carrying the host's windowed DNS behaviour

A lexical trigger and a behavioural trigger are both required, so a long odd name on its own is not a tunnel and a high query rate on its own is not a tunnel. One input vector carries both halves of the evidence, which is why the per-query path does not have to join against a separate window vector.

Trigger conditions for the DNS_TUNNEL detector
FieldValueNote
OVERSIZED_QUERY_NAMEQuery name ≥ 60 characters
HIGH_QUERY_ENTROPY≥ 3.6 bits/char with ≥ 24-character payload
HEX_ENCODED_SUBDOMAINHexadecimal-encoded label
DEEP_ENCODED_SUBDOMAINDeeply encoded label structure
UNIQUE_SUBDOMAIN_PER_QUERY≥ 80 % unique names over ≥ 20 queries
EXCESSIVE_TXT_QUERIESTXT record volume
SUSTAINED_QUERY_RATESustained rate against the host's own window
RESPONSE_LARGER_THAN_QUERY≥ 300-byte response and ratio ≥ 6

Evidence fields

  • query_length
  • subdomain_depth
  • char_entropy
  • unique_subdomain_ratio
  • txt_query_count
Limit and recorded evidence — DNS_TUNNEL

Two thresholds were recalibrated after a mixed-traffic fixture produced a false positive. The entropy trigger's minimum payload rose from 16 to 24 characters, and the response-size trigger now needs a 300-byte floor as well as the ratio — an ordinary A-record answer is already several times the length of the query name, so a ratio alone flagged normal DNS.

Recorded detail for the DNS_TUNNEL detector
FieldValueNote
RequirementOne lexical trigger AND one behavioural trigger
Vector designPer-query values plus the host's windowed DNS behaviourOne input, both halves of the evidence — no cross-vector join on the hot path.
answer_countCapped at 64 answers, each ≤ 256 charactersFree-text and list fields are length- and count-capped at ingest. No payload content is stored.

Drawn from docs/DETECTORS.md · docs/ALERT_SCHEMA.md

TLS_MALWARE

Malware in encrypted sessions

RULEmodel 20260929-1
Entity
HOST
Window
Per TLS session, handshake metadata only

No decryption, anywhere. The detector works from handshake metadata only: JA4 and JA4S fingerprints, offered ciphers, TLS version, SNI shape, certificate characteristics, and session size and timing. The artifact is a curated, hash-verified fingerprint list rather than a trained model, and it reports itself as a rule with rule evidence, which is accurate by construction.

Trigger conditions for the TLS_MALWARE detector
FieldValueNote
FINGERPRINT_REPEATED_IN_LISTThe fingerprint appears on ≥ 2 rows of the source list
JA4_AND_JA4S_AGREEClient and server fingerprints both match

Evidence fields

  • ja4
  • ja4s
  • tls_version
  • sni
  • cert_subject_cn
  • cert_issuer_cn
  • cert_validity_days
Limit and recorded evidence — TLS_MALWARE

There is no labelled JA4/JA4S corpus in this repository, so precision, recall and false-positive rate are not measurable. Promotion of the shipped artifact required an explicit acknowledgement that those three metrics are unmeasured. A match is an investigation lead, not a verdict: legitimate software and firmware share TLS libraries.

Recorded detail for the TLS_MALWARE detector
FieldValueNote
ArtifactFingerprintMatcher, input_kind fingerprints, lookup(ja4, ja4s)Built locally from a pinned 29-row AGPL-3.0 curated list. The hash of the committed CSV is verified before the matcher is constructed.
confidence_basisRULE_EVIDENCEConfidence is evidence-weighted and capped: 0.6, plus 0.2 when both sides agree, plus 0.15 when the list repeats it. Never above 0.90.
SeverityHIGH when both sides agree and the list repeats; MEDIUM otherwise
Reason codesTLS_CURATED_JA4_MATCH · MALICIOUS_JA4_FINGERPRINT · MALICIOUS_JA4S_FINGERPRINTEvery match carries them, so the basis is visible on the alert rather than inferred.
Origin-onlyClient JA4 is retained; server JA4S is ignored unless responses were observed
QUICA benign QUIC Initial demo produces Zeek QUIC/SNI/ALPN/client-JA4 metadataThat is a metadata control, not a QUIC malware classification result. QUIC malware detection is not established.

Drawn from docs/DETECTORS.md · docs/RUNTIME_MODELS.md

EXFILTRATION

Data exfiltration

STATISTICALno model
Entity
HOST
Window
300 s, internal → external flows only

Volume is the trigger and never fires alone. Backups, CI uploads and video calls move large amounts of data outbound every day, so an outbound-dominance ratio, destination rarity or a deviation from the host's own warm baseline is required on top of the volume.

Trigger conditions for the EXFILTRATION detector
FieldValueNote
Volume gatebytes_out_total_300s ≥ 10 MB
OUTBOUND_DOMINATED_TRANSFEROutbound/inbound ratio ≥ 20
RARE_DESTINATIONDestination rarity ≥ 0.9
BASELINE_DEVIATIONZ-score ≥ 4 against the host's warm baseline
ABOVE_HISTORICAL_P95Above the learned 95th percentile
OUTBOUND_FLOW_SIZE_DEVIATIONOrigin-only path: ≥ 10 MB and z ≥ 4 on a warm per-host flow-size baselineRequires 30 configured observations. Compares before updating, learns only from eligible outbound connections, and large flows emit a host vector immediately so a short transfer is not lost between emit ticks.
RARE_DESTINATION_PORTUnfamiliar service port

Evidence fields

  • bytes_out_total_300s
  • out_in_ratio
  • destination_rarity
  • *_zscore
  • *_baseline_p95
  • *_baseline_warm
Limit and recorded evidence — EXFILTRATION

Cold start is explicit. With neither a rate baseline nor a flow-size baseline warm, the result carries BASELINE_COLD so nobody reads the alert as abnormal for a host that has not been established as normal.

Recorded detail for the EXFILTRATION detector
FieldValueNote
DirectionInternal → external flows only
Confidence weightsbaseline evidence 0.35 · volume 0.25 · ratio 0.25 · rarity 0.10Baseline evidence is weighted highest, because volume alone is the weakest signal in the set.
In the alertPrior mean, observation count, observed bytes and z-scoreThe reader gets the comparison, not just the verdict.

Drawn from docs/DETECTORS.md

ANOMALY

Behavioural anomaly

STATISTICALmodel 20260928-1
Entity
HOST
Window
Host window, 1 s through 300 s

An isolation forest over six host-window features — flow rate, outbound bytes per second, SYN rate, unique destination ports, unique destination hosts and DNS query rate — fitted to benign examples. Its output is supporting evidence about how unusual a window is. It is not a new attack category and an anomaly score alone is not a finding.

Trigger conditions for the ANOMALY detector
FieldValueNote
Model scoreAbove the recorded score threshold for the fitted estimator
Features6 host-window features

Evidence fields

  • flows_per_sec_60s
  • bytes_out_per_sec_60s
  • syn_rate_5s
  • unique_dst_ports_60s
  • unique_dst_ips_60s
  • dns_queries_per_sec_60s
Limit and recorded evidence — ANOMALY

Lab-trained with a recorded lab false-positive rate of 0.066964. A candidate fitted to public benign captures was rejected outright: benign browsing volume distributions do not transfer across captures, and malicious device histories are not necessarily volume outliers. It measured a 0.233 benign false-positive rate and missed every held-out malicious window.

Recorded detail for the ANOMALY detector
FieldValueNote
EstimatorIsolationForestUnsupervised, so no supervised feature-importance file exists and none is fabricated.
Championanomaly 20260928-1, feature schema 1.2.0
Role in fusionContributes evidence and corroboration to a classified alertAnomaly is not a threat_class on its own; fusion attaches its contribution to a classified finding.

Drawn from docs/RUNTIME_MODELS.md · ANOMALY_MODEL_CARD.md

DGA and DNS tunnelling are not the same problem

Both read DNS query names, and they are easy to conflate. They are separate detectors with separate triggers, and conflating them would make both less explainable.

DGA

An algorithmically generated name looks unusual. The question is whether this single name is machine-generated — a lexical classification over 14 features of the name itself, scoring one query at a time. A supervised model does the classification because the pattern is in the string.

It says nothing about volume. One odd name is one odd name.

DNS tunnelling

A legitimate-looking name is carrying data. The question is about the volume and shape of a host's DNS behaviour over time — a lexical trigger and a behavioural trigger, both required. No model is involved; the conditions are counted.

Encoded tunnel traffic can use perfectly plausible-looking labels. Judging the name alone would miss it.

How a detector hit becomes an alert

One detector is a signal. An alert is the product of several, and the record says which.

Confidence, and what it is not

Confidence is how strongly the evidence supports the classification, and the basis is recorded on every result: rule evidence, a normalised anomaly, a calibrated probability, or fusion. Every value is reproducible from the stored supporting features.

Severity is a different question — potential impact — and is computed separately. High confidence in a small event stays low severity. There is a test asserting exactly that.

Corroboration, bounded and recomputable

Confidence is the primary value plus 0.08 per distinct corroborating detector, capped at 0.99. Bounded, monotonic, and recomputable from the stored detector contributions rather than taken on trust.

Two or more distinct threat classes on one subject within 300 seconds become an incident. One class is an alert; two is a story.

Why a challenger can never alert

Challenger results are persisted with the shadow flag set, and fusion drops them. A candidate model can be scored against live traffic for comparison, and that is the full extent of its reach.

What severity is allowed to imply

A severity floor applies to classes that imply compromise: C2 beaconing, exfiltration and DNS tunnelling start no lower than medium. Two other detectors independently agreeing escalates one step.

Severity is never carried by colour alone — every severity has a text label and a distinct marker shape.

What has been measured, and what has not

Historical single-process measurements on named hardware. The runtime has changed since; these are recorded results, not a capacity claim.

Recorded benchmark measurements and the conditions they were taken under
FieldValueNote
Mixed synthetic scenario1 584 events/secSingle process, single worker, full persistence. 121 664 events, Linux x86_64, 12 logical CPUs, Python 3.12.13.
Stage attributionFeature engine 67 %Of 76.8 s wall time: feature engine 51.5 s, persistence and generation 29 %, fusion 3 %, detector inference 0.6 %.
Per-event latencyp50 0.150 ms · p95 1.36 ms · p99 2.05 msThe 572 ms maximum is a 500-event persistence flush plus a SQLite WAL checkpoint — a throughput artefact, not detection latency.
Real Zeek replay836.956 events/sec over 100 eventsOne PCAP, one alert, p50 0.782 ms. Valid for that fixture and that hardware only. Re-measure against deployment traffic before quoting a capacity figure.
Model inference1.0–2.4 ms per rowMeasured single-row predict_proba. Cross-validated calibration measured ~7.9 ms and breached the 5 ms gate, which is why calibration is fitted prefit on a frozen estimator.

Not measured

Recorded as gaps rather than guessed. A number that was not taken is not a zero.

  • Redis round-trip throughput as a separate figure
  • Multi-worker scaling — available through consumer groups, not measured
  • Mbps — the generator produces flow metadata, not packet volume, so a line-rate figure would be fabricated
  • Memory and CPU ceilings under sustained load
  • A ClickHouse comparison — the store is SQLite in this prototype

Model quality, benchmark methodology and the full training and validation approach are in the research notes.