Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports
Reproducibility audit finds that published accuracy metrics for knowledge-graph extraction from threat reports are highly sensitive to the triple-matching method used, undermining tool comparisons.
Summary written by editorial AI · Source link below
arXiv:2609.01671v1 Announce Type: new Abstract: Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve inspected systems. Re-scoring ten system outputs on shared documents under eight protocols reverses eleven of forty-five pairwise orderings; one fixed predictio
Editorial Analysis
Security teams relying on published benchmarks to choose threat-intel KG tools may be comparing apples to oranges if matching methodologies differ.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the Research Desk
- 39 New Methods That Compromise Passkey Authentication3d
- Security Vulnerability in a Voting System3d
- Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification4d
- How Reliable Is the Multi-Input Heuristic for Bitcoin Address Clustering in Law Enforcement Contexts?4d
- Privacy Leakage in Federated Learning: Gradient-Based Client Identity Inference and Defenses for Inertial Sensing in Vehicular Edge Networks4d