C2, botnet and DGA
C2 beaconing and DGA are two different questions asked of the same traffic. The repository holds material for both, and it is where most of the honest negative results live: two calibrated DGA candidates trained on public corpora were evaluated and rejected rather than promoted.
Sources
| Field | Value | Note |
|---|---|---|
| IoT-23 | Garcia, Parmisano & Erquiaga (2020) · DOI 10.5281/zenodo.4743746 | Released under CC-BY-4.0. Individual conn.log.labeled files for malware captures 8, 20, 21, 34, 42 and 44 — 44,706 original flows. Supplied the primary supervised C2 labels, with raw labels retained and both det_label and detailed-label header variants decoded. |
| CTU-13 Neris | CTU-Malware-Capture-Botnet-43 · 36,261,479 bytes | 40,198 normalized events in transfer replay. Used separately for raw PCAP transfer evidence. Broad botnet traffic is explicitly not labelled wholesale as C2. |
| CTU Normal captures | CTU-Normal-20 · CTU-Normal-21 · CTU-Normal-22 | 282,415,864 / 311,638,284 byte benign PCAPs and the original conn/dns/ssl Bro logs for Normal-22, giving 18,892 / 10,662 / 76,166 normalized events. |
| Chrmor DGA corpus | 25 Netlab-360 DGA families plus Alexa benign names | An independent real-world corpus, used specifically because an algorithm-derived wordlist has no held-out generator to test generalisation against. 557,717 eligible rows. Three families — necurs, kraken, symmi — were held out entirely. |
| Public DGA wordlists | andrewaeva/DGA · GNU GPL v2 | Conficker, Cryptolocker, Matsnu, Pushdo, Ramdo, Rovnix, Tinba and Zeus. At most 25,000 distinct domains per family with family metadata retained. These are algorithm-derived wordlists, not observed infection telemetry. |
| Benign ranking | Tranco 64XQX · fixed date 2026-09-14 | 1,000,000 ranked domains, 200,000 distinct accepted domains before eligibility filtering. Cite Le Pochat et al. and the fixed list. Constituent terms include Majestic CC-BY-3.0, CrUX CC-BY-SA-4.0 and Radar CC-BY-NC-4.0; no blanket commercial permission is inferred. |
Recorded results
- DGA candidate on the public wordlist corpus: precision 0.974264, recall 0.739983, F1 0.841114, FPR 0.0198822, PR-AUC 0.975093, ROC-AUC 0.976351 over 281,847 eligible rows. Unseen-family recall 0.15914. Rejected — recall is below the 0.80 governance floor, with no exception requested.
- DGA candidate on the independent real-world corpus: precision 0.974928, recall 0.625296, F1 0.761916, FPR 0.017380 over 557,717 eligible rows. Unseen-family recall 0.311694. Also rejected on recall.
- The measured finding that matters: on the independent corpus the two transfer probes that condemned the first candidate both pass. That isolates the first failure as generator memorisation of a synthetic-lab corpus rather than a ceiling of the 14 lexical features.
- C2 candidate on IoT-23 labels with CTU benign captures: capture-disjoint splits of 16,275 / 14,072 / 21,997, with only 26 final positives. Every trial collapsed to one calibrated validation score; recall 0, precision 0, PR-AUC 0.001182. Rejected — the classifier detects nothing.
- A GRU challenger was not trained. PyTorch is absent and 26 positive test windows with a failed baseline do not support a credible promotion, so no GRU score or comparison is claimed.
Limit CSE-CIC-IDS2018 was not downloaded. The public run does not resolve temporal mismatch between modern benign rankings and older DGA algorithms, benign-proxy contamination, narrow device coverage, or the extreme final-class imbalance in the C2 set. A rejected candidate is still a measured result — the point of recording it is that the next attempt does not repeat it.
Drawn from docs/DATASET_PROVENANCE.md · docs/ML_EVALUATION.md · C2_MODEL_CARD.md · DGA_MODEL_CARD.md