← research
Verified research11 October 2026Checked 11 October 20267 min read

Why 32 Threat Leads Did Not Become Hunt Packages

A production snapshot of ThreatWatch shows how corroboration, observables, ATT&CK context, and readiness thresholds keep developing leads out of the qualified hunt queue.

ThreatWatchThreat intelligenceEvidence gates
Thirty-three ThreatWatch candidates separated into one qualified hunt package and thirty-two developing leads
ThreatWatch public API snapshot. Counts are time-bound and may change after a later pipeline run.

I expected the qualified list to be short. I did not expect it to be one.

On 11 October 2026, I took a production snapshot from the public ThreatWatch hunt API. It contained 33 candidates: one qualified hunt package and 32 developing leads.

33candidates reviewed
1qualified package
32developing leads
4cb1650deployed revision

At first glance, that can look like a poor conversion rate. It is actually the behaviour I want from the system. Threat reporting is abundant. Evidence that can support a bounded hunt is much rarer. The gate should stop an attractive but incomplete idea before an analyst mistakes it for finished work.

Question

What separates a qualified ThreatWatch hunt package from a developing lead, and did the production output follow that rule in this snapshot?

I am not asking whether every underlying report is true or whether the qualified item is malicious in a particular environment. Those questions require local telemetry and analyst validation. This review is narrower: whether the system’s published classification is consistent with its documented and tested evidence gate.

Method

I checked four layers that can be inspected independently:

  1. The public API response, including its generation timestamp and qualified and developing counts.
  2. The gate constants and qualification logic in modules/hunt_engine.py at deployed revision 4cb1650.
  3. The behavioural tests in tests/test_hunt_engine.py.
  4. The public hunt-package contract in docs/HUNTS.md.

I also checked the production health response before treating the API as current. It reported an OK state and deployed revision 4cb1650cbc090cb0c13a2abf6d7388e67f138614. The API response had been generated at 2026-10-11T11:10:43.827487+00:00.

That makes this a time-bound production observation tied to a specific implementation, not a claim about every future pipeline run. I preserved the aggregate values in a small public record so the numbers in this article can be checked without exposing the underlying observable or raw logs: Download the sanitised evidence record.

ThreatWatch evidence gate from public reports through corroboration, observable, ATT&CK mapping, and readiness threshold
The qualification path implemented at revision 4cb1650. A candidate that misses a required gate remains a developing lead.

The gate in code

The implementation sets three numeric thresholds: MIN_REPORTS = 2, MIN_PUBLISHERS = 2, and QUALIFICATION_SCORE = 60. A score alone is not enough. A candidate also needs an actionable observable and an ATT&CK mapping.

In practical terms, the gate asks five questions:

Gate Requirement Why it matters
Corroboration At least two reports One article can repeat an unsupported claim.
Publisher diversity At least two publishers Syndication from one origin is not independent support.
Observable At least one actionable value A hunt needs something that can be queried or matched.
Behaviour At least one ATT&CK technique The package needs adversary-behaviour context.
Readiness Score of 60 or higher Evidence quality must clear a declared threshold.

There is one deliberate exception. A structured provider match can corroborate a candidate that has only one conventional report. That is not a free pass: the observable, ATT&CK, and readiness requirements still apply. The exception is explicit in code and covered by a test, which makes it inspectable rather than an undocumented judgement.

Evidence from the snapshot

The API returned 33 candidates. One package was classified as qualified with a readiness score of 100. Its published evidence summary recorded 8 reports from 8 publishers, 1 actionable observable, 3 ATT&CK techniques, and a positive match in the CISA Known Exploited Vulnerabilities Catalog.

The other 32 candidates remained developing leads. The API does not hide them. It keeps their available evidence and limitations visible without presenting them as ready hunt packages. A developing lead can still guide monitoring, collection planning, or further research. It just has not earned operational confidence yet.

The counts also align with the code path. Candidates are evaluated against the gate, labelled with qualifications and limitations, sorted, and returned with separate qualified and lead counts. The production response exposed exactly those two groups.

What the tests prove

The tests cover the failure modes that are easy to blur in prose:

  • A multi-source candidate with an observable and ATT&CK mapping qualifies.
  • A single-report candidate remains a developing lead.
  • A structured provider match can satisfy the documented corroboration exception.
  • Publisher domains do not become actor-based hunt packages merely because they appear in the source material.

These tests do not prove that every source is accurate. They prove something more specific and useful for engineering: the gate behaves consistently for the cases the project claims to support.

What surprised me

The single qualified result was not the most interesting part. The useful part was seeing 32 unfinished ideas remain visible without being dressed up as detections.

That is harder than simply producing more output. A pipeline has to preserve enough context for later investigation while being honest about what is missing. In this snapshot, the boundary held: weak candidates were not discarded, but they were not promoted either.

I also found the explicit exception worth checking. A structured provider match can supply corroboration when only one conventional report exists, yet it cannot bypass the observable, ATT&CK, or readiness requirements. That is a narrow, testable rule rather than a hidden shortcut.

Finding

For this production snapshot, the public classification is consistent with the implementation, tests, and documented contract. Thirty-three candidates entered the evaluation path. Evidence qualified one. The remaining 32 were retained as developing leads rather than inflated into hunt packages.

That is a healthier result than a large catalogue of polished but weak detections. A security product should make uncertainty legible. Here, the useful output is not only the one package that passed. It is also the visible boundary around the 32 that did not.

Limitations

This is a snapshot, not a longitudinal study. Feed health, source availability, and candidate counts can change after the recorded generation time. Repeating the request later may produce different totals.

The review validates classification behaviour, not exploitability or compromise in a reader’s environment. A qualified ThreatWatch package is a research artefact for analyst validation. It is not a production detection, a verdict, or proof that any asset is affected.

Finally, public-source corroboration can inherit errors shared across publishers. Two reports are a minimum gate, not a guarantee of truth. Local telemetry, environment-specific scope, and human review remain necessary before operational action.