InfoGuard AG (Headquarter)
Lindenstrasse 10
6340 Baar
Switzerland
InfoGuard AG
Stauffacherstrasse 141
3014 Bern
Switzerland
InfoGuard Deutschland GmbH
Frankfurter Straße 233
63263 Neu-Isenburg
Germany
InfoGuard Deutschland GmbH
Landsberger Straße 302
80687 Munich
Germany
InfoGuard Deutschland GmbH
Am Gierath 20A
40885 Ratingen
Germany
InfoGuard GmbH
Kohlmarkt 8-10
1010 Vienna
Austria
AI now writes safety analyses. In the SOC, language models generate the initial draft. They summarize security data, identify patterns, and propose an initial assessment in seconds—tasks that would take experts significantly longer to complete. This shifts the focus. Today, the key question is how to reliably verify whether the AI analysis is correct. An incorrect assessment in asecurity analysis costs more than just time. If a real attack goes unnoticed, attackers can spread further within the network and endanger critical systems and data—a risk that deserves special attention given the current threat landscape. We analyzed this in an earlier post.
Many teams approach AI-driven analyses with full oversight: every result is reviewed by a human before it reaches the customer. While this sounds responsible, it largely negates the efficiency gains provided by AI.
We call this effect the “trust tax”: the additional verification effort that arises when teams do not sufficiently trust AI results. Analysts effectively review the analysis a second time. The time saved is lost—as is the hoped-for efficiency.
The opposite mistake is just as real, but it goes unnoticed more easily. If the AI was right a hundred times, the analytics team will nod off the hundred-and-first result without checking it. Researchers refer to this pattern as “automation bias”: an excessive reliance on automated results. This error often goes unnoticed until it’s too late.
Thus, both errors stem from the same cause: trust is felt rather than measured. Without measurement, this trust cannot be reliably calibrated.
"Trust in AI does not arise simply from believing in it. Trust arises when AI results are reviewed by experts."
For decades, the industry has ensured quality through statistical sampling, without inspecting every single product. No one tests every single screw in a shipment. A statistically determined sample identifies quality deviations and enables a well-founded decision regarding the entire lot. The risks associated with acceptance and rejection remain calculable.
This is precisely what the ISO 2859-1 standard describes. Its scope is explicitly not limited to physical products; it includes data, records, and administrative processes. Even the documented analysis of a resolved security incident can be evaluated to determine whether it meets the defined requirements. This allows the underlying inspection logic to be applied to security analyses. This requires three steps.
|
Batches. All AI-supported incidents within a customer segment over a given time window form a test batch. The batches are segmented so that they are comparable within themselves: by customer risk class, incident type, and use case. A phishing incident at a regulated financial services provider and a low-severity alert at a small manufacturing company do not belong in the same category. Error classes. Class A includes anything that can cause actual harm. This includes an overlooked or misclassified high-severity incident, as well as systematically concealed artifacts or a violation of mandatory escalation rules. Class B covers serious deficiencies without immediate security consequences: for example, an incomprehensible justification, an undocumented correction, or deviations from the playbook. Class C covers procedural errors. Tolerance Values. An AQL (Acceptable Quality Level) is established for each class—that is, a defined quality level for the maximum tolerated proportion of nonconforming units in the process average. For Class A, this value is very small; for Class C, it may be significantly larger. The sample size and rejection count are determined by the lot size and inspection level. No discretion, no discussion. Just a table. |
If an incident contains at least one Class A deviation, it is assigned to the highest defect class. The most serious finding counts. This ensures that a safety-related problem cannot be hidden behind several formal deviations.
AI analysis quality determines how closely the SOC scrutinizes
Depending on the measured analysis quality, the SOC adjusts the intensity of its inspections. If rejected lots become frequent, the system switches to more rigorous testing. The sample sizes become larger and the acceptance limits tighter—targeted specifically where quality defects occur. If the analysis quality remains high over a defined sequence of batches, the SOC reduces the inspection depth. With consistently stable quality, individual batches can be accepted at random even without a detailed inspection.
The inspection depth, which is based on analysis quality, achieves more than just efficiency. It is a calibration of trust, cast into rules.
This approach accounts for errors made by both AI and humans. That is why it tightens controls as soon as quality data indicates the need to do so. For analysts, this represents a key difference: the level of inspection is determined by predefined rules and the measured analysis quality, rather than by individual incidents or situational management decisions. This ensures that all stakeholders know under what conditions the level of control increases or decreases. Consistent quality predictably leads to reduced audit effort.
Audit results only strengthen cyber defense when they trigger concrete improvements. Every recorded deviation is therefore directed back to where it can have an impact:
Governance: Unclear guidelines or escalation paths are identified and adjusted.
Training and Coaching: Recurring behavioral patterns reveal where analysts need targeted support.
Knowledge Base: Individual experiences are transformed into structured knowledge about typical AI error patterns.
Technical Optimization: Insights are incorporated into prompt design, detection engineering, and rule logic.
Over time, this creates a structured quality dataset. It reveals where collaboration between humans and AI actually breaks down and which patterns recur. It also provides insights into which incident types are suitable for greater automation and where processes, playbooks, or rule logic need to be adjusted.
A quality assurance process must function within 24/7 SOC operations. That is why we have transferred the sampling logic into a separate QA application, which is currently running in production at the Cyber Defense Center (CDC).
This app controls the selection, review, and documentation of cases. Filter groups categorize cases by sensor type and criticality. Each group has its own review quota. Critical cases are selected for quality review significantly more often than cases with lower criticality. Selection occurs automatically and randomly on a per-calendar-week basis. This ensures that there is no influence over which cases are reviewed. The selected cases are placed in a review queue and examined there using a mandatory checklist.
A dashboard provides transparency regarding the quality review at all times. It displays the number of cases analyzed, the percentage reviewed, the findings by filter group, the rejection rate, and the review effort. A sampling history documents the review rate for each week. Without this traceability, the entire process would be worthless.
AI in the SOC only delivers a security benefit if the quality of its results can be reliably verified. Measurable quality data provides a robust foundation for this:
Detecting errors: Misprioritizations become apparent earlier.
Examine risks more selectively: Critical cases receive more in-depth scrutiny.
Gain efficiency: Consistent analysis quality reduces unnecessary audit effort.
AI is transforming security analysis faster than governance models can keep up. Policies and human review alone are not sufficient. Measurable quality criteria are crucial: statistical sampling, defined tolerance limits, and binding rules for the depth of review.
Those who use AI-powered cyber defense responsibly ensure that speed does not come at the expense of reliability. Security decisions remain transparent and robust even as automation increases.
When a real cyberattack occurs, its detection must not depend on trust in an AI result. Our Management Detection and Response (MDR) team therefore combines AI-powered analytics with measurable quality assurance and assessment by experienced analysts. Contact us at to learn how our SOC and MDR services can strengthen your cyber defenses.
Caption: AI-generated image