Cybersecurity training & evaluation index · 2026
Cybersecurity RL environments, training data & evaluations.
An index of companies building cybersecurity RL environments, training data, benchmarks and evaluations for advanced AI systems.
Leaderboard · September 2026
Cybersecurity AI training and evaluation providers
29 ranked providers · 13 monitored (NR). Expand an entry for its role, evidence and sources. Equal assessments share a rank.
01Gray Swan AI↗Builds agent-security evaluations, protective methods and adversarial training systems, with substantial joint cyber research.Offensive securityAgent securityIPI; AgentHarm; ARTEMIS; Circuit Breakers; ShadePublished evidenceCommercialPittsburgh, USCorroborated
Builds agent-security evaluations, protective methods and adversarial training systems, with substantial joint cyber research.
Published agent and enterprise evaluations, protective tests and implemented attack optimization. Independent model assessments corroborate use. Joint research retains its collaborators’ credit; access varies by artifact.
02Bugcrowd↗Contributes executable exploitation evaluations and offers cybersecurity RL environments.Vulnerability discoveryOffensive securitySecure coding & patchingExploitBenchPublished evidenceCommercialSan Francisco, USCorroborated
Contributes executable exploitation evaluations and offers cybersecurity RL environments.
ExploitBench has published methods and results and independent execution reported by Anthropic, OpenAI and government evaluators. The separate advertised training curriculum is not treated as five completed benchmarks. Its separate RL Environments product supports exploit, discovery and patching objectives; the launch alone does not establish additional completed outcomes or held-out training gains.
03Dreadnode↗Combines AI-security attack benchmarks, live intrusion investigation evaluations and reusable training trajectories.Offensive securityDefensive securityTraining dataAIRTBench; DreadIndex; Ares/DreadGOAD; WorldsPublished evidenceCommercialNot listedCorroborated
Combines AI-security attack benchmarks, live intrusion investigation evaluations and reusable training trajectories.
Published attack and investigation results, an implemented AMSI blocking study, scored traces and synthetic-training experiments. GovTech and participant writeups corroborate delivery and use of its AI-security challenges. Completed evaluations and training assets; access varies by artifact, with some tasks requiring platform access. Some DreadIndex results retain flagged cheating; its separate audit distinguishes clean solves. Clean final held-out training gains remain unverified.
- AIRTBench paper ↗
- AIRTBench harness and platform dependency ↗
- Released AIRTBench scored trajectories ↗
- Implemented AMSI blocking evaluation ↗
- Ares and DreadGOAD completed offense/investigation evaluation ↗
- DreadIndex methods and score limitations ↗
- Controlled cheating and adjudication study ↗
- Worlds training experiment and limitations ↗
- Implemented Worlds training integration ↗
- GovTech confirms AI CTF vendor ↗
- Independent team's completed challenge writeups ↗
04Irregular↗Tests offensive cyber agents on realistic systems and constrained multistage objectives.Offensive securityVulnerability discoveryRL environmentsFrontierCyber; CyScenarioBench; Atomic Challenges; AISI advanced tasksPublished evidenceCommercialTel Aviv, IsraelCorroborated
Tests offensive cyber agents on realistic systems and constrained multistage objectives.
Completed provider and outside evaluations include AISI advanced tasks and Anthropic use of a CyScenarioBench subset. AISI’s wider suite and Crystal Peak’s rust_vm task are not attributed solely to Irregular. Task access varies.
05Virtue AI↗Evaluates vulnerability diagnosis, secure coding and repair, plus dynamically checked offensive tasks.Secure coding & patchingVulnerability discoveryOffensive securitySeCodePLTPublished evidenceCommercialNot listedCorroborated
Evaluates vulnerability diagnosis, secure coding and repair, plus dynamically checked offensive tasks.
Public paper, data and code; independently reused in MT-Sec. Some coding checks are rule-based, and simulated attacks do not establish real-world operational success. Released benchmark; completed independent reuse.
- Version-specific SeCodePLT methods, attribution and Appendix G ↗
- Company research catalog ↗
- Released evaluation and dataset-generation code ↗
- UK AISI counterpart confirmation ↗
- Independent MT-Sec reuse and evaluation ↗
- Implemented SeCodePLT training integration ↗
- Shared SeCodePLT/OpenSage author and Virtue affiliation ↗
06Hack The Box↗Supplies evaluated cyber ranges and tasks, with its own published agent benchmark.Offensive securityRL environmentsCooling Tower; HTB OWASP benchmarkPublished evidenceCommercialNot listedCorroborated
Supplies evaluated cyber ranges and tasks, with its own published agent benchmark.
AISI evaluates Cooling Tower. HTB publishes OWASP task results and grading methods; Google DeepMind and D-CIPHER independently evaluate other HTB tasks. Completed evaluations and commercial AI Range; access varies by asset.
- AISI range methods, attribution and per-step results (March 2026) ↗
- HTB OWASP benchmark: task results, fresh instances and validated-flag grading ↗
- Google DeepMind independent HTB evaluation, section 4 and acknowledgments ↗
- D-CIPHER independent HTB execution: sections 5–6 and Table 3 ↗
- HTB identifies Cooling Tower as its range (May 20, 2026) ↗
- HTB AI Range offering; commercial context rather than training proof ↗
- Implemented CTF MCP access; evaluation integration ↗
07Collinear AI↗Evaluates secure code repair with functional and security checks, alongside a separate cyber training offering.Secure coding & patchingDefensive securityRL environmentsCWE-BenchPublished evidenceCommercialSan Francisco, USCorroborated
Evaluates secure code repair with functional and security checks, alongside a separate cyber training offering.
CWE-Bench publishes methods and task-linked results; Google reports a completed Gemini evaluation on the benchmark. Tasks are available through qualified access.
08Trajectory Labs↗Supplies prompt-injection and coding-agent security evaluations documented by model developers.Agent securityOffensive securitySearch-PI Complex; coding-agent security evaluationsPublished evidenceCommercialBerkeley / Remote, USCorroborated
Supplies prompt-injection and coding-agent security evaluations documented by model developers.
Meta and Anthropic report specific completed evaluations and their conditions. These are counterpart-confirmed evaluations; public task distribution is not required.
09Endor Labs↗Evaluates secure code changes and vulnerability finding, with detailed benchmark-integrity analysis.Secure coding & patchingVulnerability discoveryAgent Security League; AI-SAST evaluationPublished evidenceCommercialNot listedDocumented
Evaluates secure code changes and vulnerability finding, with detailed benchmark-integrity analysis.
Measured results and integrity controls are public. Independent replication of Endor's evaluations remains unverified. Completed public evaluations.
=10Incalmo↗Builds cyber-agent environments and toolkits with completed range and malware evaluations.Offensive securityRL environmentsCyber ranges; PathoGenPublished evidenceCommercialUndisclosedCorroborated
Builds cyber-agent environments and toolkits with completed range and malware evaluations.
Anthropic reports toolkit and range evaluations; Incalmo publishes attributable range and PathoGen results. These support technical use under documented conditions, with task access varying by asset. Academic toolkit and range results are distinguished from commercial-product performance claims.
=10LogicStar↗BaxBench tests generated backend applications for functional correctness and resistance to executed security exploits.Secure coding & patchingTraining dataBaxBenchPublished evidenceCommercialNot listedCorroborated
BaxBench tests generated backend applications for functional correctness and resistance to executed security exploits.
Public tasks, code and measured results; independently extended and evaluated in MT-Sec. Secure construction is counted once. Released benchmark; completed independent reuse.
=10SpecterOps↗Built the enterprise-intrusion range used in AISI frontier-model evaluations.Offensive securityRL environmentsThe Last Ones — AISI enterprise cyber rangePublished evidenceCommercialNot listedCorroborated
Built the enterprise-intrusion range used in AISI frontier-model evaluations.
AISI publishes methods and measured progress through the range. SpecterOps separately confirms its construction role. Completed external evaluations; public methods and results, restricted range access.
13Invariant Labs↗Co-developed a runnable framework for evaluating prompt injection and protective interventions in tool-using agents.Agent securityAgentDojoPublished evidenceCommercial + OSSZürich, SwitzerlandCorroborated
Co-developed a runnable framework for evaluating prompt injection and protective interventions in tool-using agents.
Public code, attack and utility checks, and measured defenses support the evaluation. Scale’s ASPI work provides outside reuse. Credit is limited to the documented joint contribution.
=14HUD↗Co-developed ZeroDayBench and publishes a runnable cyber patching environment.Secure coding & patchingRL environmentsZeroDayBench; MLflow patching environmentPublished evidenceCommercialNot listedDocumented
Co-developed ZeroDayBench and publishes a runnable cyber patching environment.
ZeroDayBench reports patching results across five information levels in real repositories with deliberately inserted vulnerabilities. HUD is a joint contributor and publishes a runnable MLflow example with outcome checks. The example is not the full suite; independent benchmark use and controlled cyber training gains were not established.
=14MatterSec Labs↗Measures vulnerability diagnosis in code and repositories, with results broken down by security category and stakeholder priorities.Vulnerability discoveryDefensive securitySecLens / SecLens-RPublished evidenceCommercialNot listedDocumented
Measures vulnerability diagnosis in code and repositories, with results broken down by security category and stakeholder priorities.
Released evaluation and scoring code, paired vulnerable/patched cases and measured category results. A Kalmantic coauthor confirms the collaboration; independent execution remains unverified. Public code, methods and results; completed joint research.
=14Simbian↗Evaluates cyber investigations using attack telemetry and machine-checkable outcomes.Defensive securityTraining dataRL environmentsCyber Defense BenchmarkPublished evidenceCommercialMountain View, USDocumented
Evaluates cyber investigations using attack telemetry and machine-checkable outcomes.
Published benchmark code and results support the defense evaluation; separate training assets support training utility. Provider-run model comparisons do not establish independent adoption or controlled learning gains. Public harness and sample; full benchmark dataset available by request.
17Turing↗Evaluates cyber agents across defensive, offensive and forensic tasks.Defensive securityOffensive securityRL environmentsCyberStrikePublished evidenceCommercialPalo Alto, USDocumented
Evaluates cyber agents across defensive, offensive and forensic tasks.
Published task contracts, grading methods and measured results support the released evaluation. Task access is gated, and independent outside execution remains unverified. Counts vary between the article and dataset card, so no combined count is asserted.
18Scale AI↗Evaluates agent prompt injection and protective interventions, with a separate cyber-relevant propensity track.Agent securityOffensive securityASPI; PropensityBench cyber trackPublished evidenceCommercialSan Francisco, USDocumented
Evaluates agent prompt injection and protective interventions, with a separate cyber-relevant propensity track.
Public methods, implementations and measured outcomes support agent-security work. General model-provider relationships and unrelated PropensityBench tracks do not establish additional cyber validation. PropensityBench measures unsafe choices using simulated proxy tools; cyber is one of four domains, rather than a measured offensive capability.
=19ARIMLABS↗Evaluates malware reverse engineering using executable samples and checked investigation outputs.Defensive securityTraining dataMalware BenchPublished evidenceResearch labWarsaw, PolandDocumented
Evaluates malware reverse engineering using executable samples and checked investigation outputs.
Published tasks, scoring, results and traces support the benchmark and training utility. Independent outside execution remains unverified; no nonprofit status is assumed.
=19Quesma↗Measures agents’ ability to find deliberately introduced backdoors in real software binaries.Vulnerability discoveryDefensive securityBinaryAuditPublished evidenceCommercial + OSSWarsaw, PolandDocumented
Measures agents’ ability to find deliberately introduced backdoors in real software binaries.
Published task implementations, grading methods and task-linked results establish a diagnostic benchmark. Finding backdoors is not counted as a separately demonstrated exploitation or repair capability.
=21Vals AI↗Evaluates vulnerability reproduction and patching with distinct task contracts and measured results.Secure coding & patchingVulnerability discoveryCyberBenchPublished evidenceCommercialSan Francisco, USDocumented
Evaluates vulnerability reproduction and patching with distinct task contracts and measured results.
Public methods, track results and patch examples support completed work. Restricted tasks are eligible; independent outside execution remains unverified in this review. CyberBench builds on CyberGym methodology, ARVO reproduction images and mini-swe-agent. Results are public; the task set is held out.
=21Vercel↗Evaluates model vulnerability findings and releases the agent harness used for repository security reviews.Vulnerability discoveryDefensive securityDeepsecBench / deepsecPublished evidenceCommercialNot listedCorroborated
Evaluates model vulnerability findings and releases the agent harness used for repository security reviews.
Published benchmark methods and model results; Roboto Studio documents a completed independent deepsec run on its own application, including configuration and outcome limitations. Public harness and benchmark results; benchmark repository and reference findings remain confidential.
23Vmax↗Studies procedural Unix security tasks and reinforcement learning with controlled evaluation.Offensive securityRL environmentsunix-ctfPublished evidenceCommercialSan Francisco, USDocumented
Studies procedural Unix security tasks and reinforcement learning with controlled evaluation.
Published methods report training gains with held-out and baseline controls. The implementation is not publicly released; that does not reduce eligibility. Single-seed reporting limits conclusions about stability.
24Snorkel↗Contributes executable security tasks within broader agent evaluations.Defensive securityRL environmentsShadowRelay; OpenThoughts-TBLite security tasksPublished evidenceCommercial + OSSRedwood City, USCorroborated
Contributes executable security tasks within broader agent evaluations.
Named security work and joint OpenThoughts-TBLite tasks have published tests and results, with outside integration documented by Nous Research. Credit applies to the attributable cyber contribution, not every task in the wider suite.
25Crystal Peak Security↗Contributed advanced cyber tasks used in AISI model evaluations.Offensive securityAISI advanced cyber tasks — rust_vm contributionPublished evidenceCommercialNot listedCorroborated
Contributed advanced cyber tasks used in AISI model evaluations.
AISI explicitly attributes rust_vm to Crystal Peak and publishes its task mechanics, intermediate checks and completed model result. The wider suite is joint work with Irregular. Completed externally reported evaluation; public methods and results, task access not established.
26Prime Intellect↗Provides infrastructure used by third-party cyber evaluation and training projects.Agent securityRL environmentsTraining dataSecurity Verifiers; hosted Strix-XSS trainingPublished evidenceCommercial + OSSSan Francisco, USCorroborated
Provides infrastructure used by third-party cyber evaluation and training projects.
Security Verifiers documents integration and release boundaries; a Strix-XSS model record attributes hosted training. Prime receives infrastructure credit, not authorship of those third-party benchmarks.
27Bespoke Labs↗Contributes to a joint agent-task release that includes executable security work.Secure coding & patchingRL environmentsOpenThoughts-TBLite security tasksPublished evidenceCommercial + OSSMountain View, USCorroborated
Contributes to a joint agent-task release that includes executable security work.
Public task definitions, tests and aggregate results, plus Nous Research integration, support the contribution. Broader suite ownership and training gains are not inferred from joint participation.
28AfterQuery↗Measures vulnerability analysis on real code with human-graded findings.Vulnerability discoveryTraining dataVADERPublished evidenceCommercial + OSSSan Francisco, USDocumented
Measures vulnerability analysis on real code with human-graded findings.
The paper, case data and graded outcomes document a completed diagnostic evaluation. Static data and human grading are distinguished from an interactive training environment.
29Labelbox↗Evaluates privacy and security behavior in bounded agent simulations.Agent securityTraining dataRL environmentsImplicit Intelligence — Privacy & SecurityPublished evidenceCommercialSan Francisco, USDocumented
Evaluates privacy and security behavior in bounded agent simulations.
The named Privacy & Security track supplies cyber-relevant tasks and measured results. General harmful-content jailbreak research is background and does not establish additional cyber outcomes.
NRDatacurve↗Cybersecurity RL product category is explicit, but a concrete qualifying cyber artifact or delivered evaluation is missing.Secure coding & patchingTraining dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited
Cybersecurity RL product category is explicit, but a concrete qualifying cyber artifact or delivered evaluation is missing.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRFleet AI↗General environments and repository forks; no attributable integrated cyber artifact or completed cyber evaluation identified.RL environmentsTraining dataQualifying evidence not verifiedEvidence gapCommercial + OSSNew York, USLimited
General environments and repository forks; no attributable integrated cyber artifact or completed cyber evaluation identified.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRHandshake AI↗Cyber recruiting is verified; attribution of a personal SOC study to a Handshake-owned platform is unresolved, and no material integrated technical role is established.Training dataOffensive securityQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited
Cyber recruiting is verified; attribution of a personal SOC study to a Handshake-owned platform is unresolved, and no material integrated technical role is established.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRHuzzle Labs↗A named access-review example without substantive task details or measured security results.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialLondon, UKLimited
A named access-review example without substantive task details or measured security results.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRIndium AI Labs↗CrashDiag is infrastructure repair; a completed substantive cybersecurity evaluation for IntrusionWatch was not verified.Defensive securityOffensive securityRL environmentsQualifying evidence not verifiedEvidence gapResearch labDistributedLimited
CrashDiag is infrastructure repair; a completed substantive cybersecurity evaluation for IntrusionWatch was not verified.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRInvisible Technologies↗Agent-security service and human expertise assessment are documented, but no qualifying completed model cyber evaluation was inspected.Training dataRL environmentsOffensive securityQualifying evidence not verifiedEvidence gapCommercialDistributedLimited
Agent-security service and human expertise assessment are documented, but no qualifying completed model cyber evaluation was inspected.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRMechanize↗General software evaluation is documented, but no qualifying cybersecurity artifact or completed evaluation.Secure coding & patchingRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited
General software evaluation is documented, but no qualifying cybersecurity artifact or completed evaluation.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRMercor↗Cyber recruiting/program descriptions do not yet establish a completed technical asset or evaluation.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited
Cyber recruiting/program descriptions do not yet establish a completed technical asset or evaluation.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRPreference Model↗Specific cyber recruiting/build plan; no completed qualifying artifact or evaluation verified.Secure coding & patchingVulnerability discoveryOffensive securityQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited
Specific cyber recruiting/build plan; no completed qualifying artifact or evaluation verified.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRRise Data Labs↗Relevant security services are advertised, but no qualifying technical artifact or completed evaluation was inspected.Agent securityRL environmentsTraining dataQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited
Relevant security services are advertised, but no qualifying technical artifact or completed evaluation was inspected.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRRL Labs↗Cyber environment descriptions without inspectable named tasks or completed measured evaluation.Offensive securityDefensive securityRL environmentsQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited
Cyber environment descriptions without inspectable named tasks or completed measured evaluation.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRSurge AI↗Broad safety collaboration and cyber-related recruiting; no attributable completed cyber-specific technical evaluation verified.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited
Broad safety collaboration and cyber-related recruiting; no attributable completed cyber-specific technical evaluation verified.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
NRVLNO↗Named partnership and claimed benchmark, but no directly inspected technical task set or completed measured evaluation.Offensive securityAgent securityRL environmentsQualifying evidence not verifiedEvidence gapCommercialTel Aviv, IsraelLimited
Named partnership and claimed benchmark, but no directly inspected technical task set or completed measured evaluation.
NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.
Methodology
Rankings combine five factors, applied consistently to completed, attributable cyber work. A provider must have a substantive technical artifact or delivered evaluation with methods and outcomes to qualify.
We assess each provider’s strongest substantial qualifying contribution for each factor. That does not imply every capability is available in one product. Tied assessments share a rank.
Open, qualified and confidential task access are treated equally when sufficient evidence can be verified. Tenure, company size, prestige and generic customer logos are not scoring factors. Roadmap announcements receive no credit for completed work.
Scope recognizes distinct outcomes but does not require offensive work. Specialists can improve through depth, realism, verification, training utility and outside validation.
Evidence labels summarize the record: Documented means qualifying technical evidence; Corroborated adds confirmed outside execution or use; Limited marks a remaining eligibility gap. These labels are not an additional scoring factor.
- 01
Technical substance, depth & scope
Completed, attributable cyber work with substantive methods and outcomes. Depth requires measured differences across mechanisms, conditions or skill stages. Scope recognizes separately evaluated diagnosis, repair/hardening, offensive effects, detection/blocking and containment/recovery. Multiple steps toward one outcome count once.
- 02
Technical realism
Real code, binaries, telemetry and operational constraints; meaningful tool use, multistep dependencies and interaction with real defenses. Task count alone does not establish realism.
- 03
Verification rigor
Checks that establish the intended security outcome, relevant negative or regression controls, repeatable evaluation and evidence addressing grader errors, shortcuts, leakage and contamination.
- 04
Training utility
A progression from static assets to repeatable evaluation, additional usable training assets or implemented training integration, and controlled gains on held-out tasks. Model scores alone do not prove training value.
- 05
External validation
Distinguish provider claims, confirmed collaboration, completed outside technical use and broader independent execution under different conditions. Repeated citations, logos and publications of the same evaluation do not create new validation.
Cybersecurity AI training FAQ
Technical answers about training and evaluating cyber agents in interactive, verifiable environments.
What is a cybersecurity RL environment?+
A cybersecurity reinforcement learning environment is an interactive system in which an AI agent performs security tasks and receives machine-checkable feedback. Unlike a static dataset, it changes in response to the agent’s actions and can measure successful outcomes across a multistep workflow.
How is a cyber RL environment different from a cyber range or CTF?+
A cyber RL environment is instrumented for repeated model training and evaluation, with reliable resets, rewards and verifiers. A cyber range primarily recreates operational infrastructure, while a capture-the-flag challenge defines a security objective. The same system can serve all three purposes when it combines realistic targets with machine-checkable outcomes.
What training data is needed to improve a secure-coding agent?+
Secure-coding agents improve on varied examples of vulnerability discovery, reproduction and patching in realistic codebases. Useful training data spans languages and weakness classes, includes failed approaches, and pairs each task with tests or programmatic verifiers that check both security and functional correctness.
How do AI agents learn vulnerability discovery and patching?+
AI agents learn vulnerability discovery and patching by working through complete security workflows: inspect a target, identify a weakness, reproduce it, produce a fix and verify the result. Interactive environments add tool use, branching paths and error recovery; held-out tasks test whether those capabilities transfer to unfamiliar code and systems.
What makes a cybersecurity task reliably verifiable?+
A reliably verifiable cybersecurity task has an objective outcome that can be checked against system state. Examples include a patch that passes functional and security tests, proof that a vulnerability was exploited, a restored service, a contained attack or a correct configuration change. Verifiers should reward the outcome rather than a particular phrase or action sequence.
How can teams prevent reward hacking and evaluation contamination?+
Teams reduce reward hacking and evaluation contamination by separating training from held-out tasks, hiding critical tests, checking final system state and looking for shortcuts that satisfy a grader without completing the objective. Strong evaluation sets are access-controlled, refreshed over time and reviewed for overlap with public data and training environments.
When should an AI lab build versus buy cyber training environments?+
An AI lab should build when the environment encodes a proprietary capability or must be tightly coupled to internal infrastructure. Buying is useful when the lab needs faster task production, specialist security expertise, broader coverage or independent held-out evaluation. Many programs combine an internal harness with externally supplied tasks, environments and audits.
How should an AI lab evaluate a cybersecurity environment provider?+
An AI lab should evaluate technical realism, verifier quality, task diversity, isolation, reset reliability, difficulty calibration and resistance to shortcuts. It should also examine whether the provider can integrate with the lab’s training stack, separate training from held-out evaluation and deliver within a model-training cycle. A public benchmark is useful evidence, but not sufficient by itself.