Image: cdn.prod.website-files.com · rights & removal
When Detection Engineering Becomes a Data Engineering Problem
Reporting by Mitiga ResearchRead the original at mitiga.io
Executive Summary
Detection engineering must incorporate data engineering principles because the scale of modern security telemetry demands it. The core problem is that detection engineers are not trained in distributed query optimization, leading to performance issues when scaling detection coverage across large data lakes. This gap manifests as detectors failing silently, as a rule exceeding its execution budget produces no alerts, and runtime becomes an unmeasured business cost. The solution involved moving from optimizing individual detectors to understanding patterns across the entire catalog, focusing on structuring the distributed query operations themselves.
The process of optimization required establishing provable safety guarantees: changes must maintain identical output for any input while reducing runtime. This led to creating a rulebook of specific, conditional optimizations that were proven safe by measuring performance against identical execution environments. The final solution involved an agentic pipeline that automates this evaluation through layered gates—cost, behavior, and filter integrity checks—ensuring that performance gains do not compromise detection accuracy or analyst experience.
Facts Only
* Execution time is a coverage concern because silence from a detector is indistinguishable from a clean environment.
* The skills gap stems from training detection engineers to find attacks rather than tune distributed queries.
* The required bar for optimization is "faster and provably identical" regarding findings, grouping, and analyst output.
* Standard Spark tuning advice can be detrimental at platform scale, as optimizations for single jobs may hurt concurrent system performance.
* A process of manual work led to a rulebook of 34 changes that upgraded over 300 detectors, with each improvement cutting runtime by 15% to 60%.
* Detection failures occur when a rule exceeds its execution budget, resulting in no output and no error signal.
* Runtime becomes a financial metric because detector runtime scales across cluster time and environment costs.
* The focus shifted from individual detectors to patterns across the catalog concerning PySpark operations like transformations.
* Cost concentration occurred disproportionately on reading CloudTrail data.
* Detail resides in requestParameters and responseElements, whose size and dynamism create parsing bottlenecks that scale with event type.
* Optimization required determining if a change produces exactly the same rows and column values for any possible input.
* A specialized pipeline was developed to automate this process through agentic execution gates.
Full Take
The narrative details a systemic failure where operational knowledge in data engineering is missing from a critical security function, leading to invisible performance and coverage deficits at scale. The pivot from treating detectors as isolated entities to treating them as elements within a larger distributed system forces an understanding of cost not just in terms of computation but in terms of observable system behavior. The pattern observed is that complexity introduced by data ingestion (dynamic JSON structures) creates hidden costs that standard single-job optimizations fail to account for.
The most crucial implication is the necessity of formalizing judgment. The realization that "use fromjson" or dropping a `distinct()` operation are not subjective preferences but verifiable, conditional statements about system invariants—safe only under specific structural conditions—is a profound shift in how expertise is codified. The agentic pipeline itself models a form of distributed reasoning: it doesn't just execute instructions; it tests hypotheses (via gates) and validates the consequences against multiple perspectives (cost, behavior, safety). This reflects a move toward systemic resilience where explicit guardrails supersede assumed best practices.
The challenge lies in ensuring that this advanced optimization is decoupled from the immediate threat of data ingestion. If the pipeline’s performance itself becomes the bottleneck, it risks becoming another layer of unmanageable complexity rather than an abstraction of necessary constraints. The question then becomes: how do we maintain cognitive sovereignty when tools are built to enforce rules, and what form does this new distributed expertise take for the detection engineer?
From the original · Mitiga Research
Key takeaways - Execution time is a coverage concern. A detector that doesn't finish produces silence, and silence is indistinguishable from a clean environment. - The skills gap is the root cause.Read the full story at mitiga.io
Sentinel — Human
This text exhibits strong signs of being human-written, characterized by deep technical expertise, idiosyncratic emphasis on process safety, and a narrative structure common in expert analysis rather than generic AI synthesis.
