As organizations increasingly employ AI in their security operations centers (SOCs) to automate processes like threat triage and incident reporting, a pressing question arises: does the effectiveness of these SOCs rely more on the selected large language model (LLM) or on the quality of the data it processes? While academic studies present mixed answers, recent findings lean strongly toward data quality being the pivotal factor in successful security operations.
A study published in Frontiers in Artificial Intelligence underscores variability in AI model outputs, citing significant differences in accuracy and clarity among leading LLMs such as Claude 3.5 Sonnet, ChatGPT 4o, and Mistral Large 2. Concurrently, research in Information Systems emphasizes that reliable AI systems necessitate high-quality training data to perform effectively.
A recent experiment from the Provably Better Data project has provided quantitative evidence supporting the argument for prioritizing data quality. The research illustrates that superior network data significantly boosts security outcomes by two to four times across critical investigation metrics.
Operational Insights for CISOs
These findings present clear advantages for Chief Information Security Officers (CISOs). They indicate a path to:
- Accelerate incident response times through improved telemetry, thereby minimizing mean time to respond (MTTR).
- Optimize resource expenditures, aiding in cost management.
- Deliver higher returns on security investments, enhancing the credibility of security teams’ contributions.
Security professionals can use these insights to inform their SOC architecture decisions, enabling infrastructure choices that mitigate analyst turnover due to alert fatigue while providing easily understandable metrics for executive teams and board members.
Structure of Research Testing
To effectively measure the factors affecting AI-driven performance within enterprise security contexts, the research incorporated a controlled testing framework evaluating model performance on two operational tasks:
- A Capture the Flag (CTF) scenario based on a Volt Typhoon attack.
- An incident response evaluation utilizing a Salt Typhoon dataset.
This framework precisely isolated data quality by processing four distinct telemetry sources under uniform conditions:
- Corelight enriched logs.
- Open-source nDPI firewall logs.
- Snort 3 intrusion detection alerts.
- NetFlow connection telemetry.
To ensure robustness and repeatability, each dataset underwent multiple tests using the Open Cybersecurity Schema Framework (OCSF). Three LLMs—Anthropic Claude Opus 4.6, Google Gemini Pro 3.1 Preview, and older iterations—were consistently tested with identical investigative prompts across all test runs, evaluating their performance based on CTF accuracy and incident response claims supported by the evidence at hand.
Performance Insights: The Boundaries of Evidence
Despite the advanced reasoning capacities of modern language models, their conclusions remain limited by the quality and depth of available data. When critical telemetry lacks full protocol-level context, AI models cannot make interpretations based on what was never recorded. The investigation process is ultimately constrained by the parameters of the data collected.
For instance, a specific question regarding the NetBIOS name for a given IP showcased striking differences based on the data sources:
- Corelight logs correctly identified the name as FINANCE01, leveraging precise NTLM log details.
- Firewall logs recognized NTLM activity but fell short in extracting the necessary details to provide an answer.
Such data deficiencies can generate notable operational challenges. Firewalls may offer broad visibility, yet missing specific protocol insights leave investigations incomplete, necessitating time-consuming manual verification by analysts. Conversely, robust data enables AI to produce actionable responses rapidly.
Quantifying the Impact of Data Quality
The measurements derived from this study vividly illustrate that high-quality data significantly enhances security outcomes, with leading datasets yielding performance improvements between two to four times compared to standard logs.
During the CTF benchmark utilizing 44 questions, accuracy rates varied markedly across the different data sources:
- Corelight logs achieved a 95.2% accuracy rate.
- Firewall logs recorded only 58.3% accuracy.
- Snort 3 alerts displayed a 39.4% accuracy rate.
- NetFlow logs lagged behind with a mere 25.8% accuracy.
Clearly, access to richer telemetry had a direct effect on how effectively the models could perform. Models powered by Corelight data achieved a CTF score of 4,178.3 points, dwarfing the 970.0 points attained using NetFlow information, marking a considerable improvement.
Similar patterns emerged in the incident response assessment, where differences in dataset coverage and total scores highlighted the crucial role of quality:
- Corelight logs provided 90.3% evidence coverage, far exceeding firewall logs at 61.3% and netting only 30.9% for NetFlow.
- Snort 3 alerts accounted for a scant 21.2% coverage.
Corelight logs allowed models to answer nearly all critical investigation questions at a rate of 91.7%, significantly surpassing the 18.3% supported by NetFlow logs and the 10% from Snort 3 alerts.
The bottom line: High-quality data generated a fivefold increase in the visibility of critical incidents in comparison to basic flow records.
Beyond accuracy, the quality of data drastically reduced investigation times. With Corelight logs, the LLM required only 14.7 minutes to complete analyses. In contrast, investigations with NetFlow logs took 27 minutes, and those with firewall logs dragged on for 26.3 minutes.
The conclusion: Inferior data significantly prolongs investigation timelines, causing models to cycle through repeated attempts.
Interestingly, the occurrence of ‘hallucinated’ outputs dropped to zero across Corelight and other datasets when models were guided to acknowledge gaps in available evidence. This contributes to reducing unnecessary investigative efforts and enhancing the confidence analysts can place in the conclusions drawn.
Strategic Takeaways for Security Leaders
AI application in security processes can significantly enhance threat detection, monitor indicators, and strengthen incident responses within enterprise networks. However, the path to effective automation is contingent upon a foundation of comprehensive, structured, and protocol-aware data.
As suggested by the findings, improvements in models and complex configurations will not compensate for intrinsic data shortcomings. Security leaders mapping out future SOC investments must give precedence to the quality of evidence over the model itself.
For an extensive breakdown of their methodology, architecture, and experimental metrics, see the Provably Better data white paper available on Corelight’s website.
Corelight Network Detection and Response
Unlocking the potential of better data can elevate security outcomes significantly. Discover why high-fidelity network evidence is vital for AI-enhanced security. Learn more.