Cybersecurity training & evaluation index · 2026

Cybersecurity RL environments, training data & evaluations.

An index of companies building cybersecurity RL environments, training data, benchmarks and evaluations for advanced AI systems.

Leaderboard · September 2026

Cybersecurity AI training and evaluation providers

29 ranked providers · 13 monitored (NR). Expand an entry for its role, evidence and sources. Equal assessments share a rank.

01Gray Swan AIBuilds agent-security evaluations, protective methods and adversarial training systems, with substantial joint cyber research.Offensive securityAgent securityIPI; AgentHarm; ARTEMIS; Circuit Breakers; ShadePublished evidenceCommercialPittsburgh, USCorroborated

Builds agent-security evaluations, protective methods and adversarial training systems, with substantial joint cyber research.

Published agent and enterprise evaluations, protective tests and implemented attack optimization. Independent model assessments corroborate use. Joint research retains its collaborators’ credit; access varies by artifact.

Offensive securityAgent security
Cyber evidence
IPI; AgentHarm; ARTEMIS; Circuit Breakers; Shade · Published evidence
Attributed role
Joint research contributor and adversarial evaluation developer
Provider type
Commercial
Location
Pittsburgh, US
Evidence
Corroborated
Visit company website ↗
02BugcrowdContributes executable exploitation evaluations and offers cybersecurity RL environments.Vulnerability discoveryOffensive securitySecure coding & patchingExploitBenchPublished evidenceCommercialSan Francisco, USCorroborated

Contributes executable exploitation evaluations and offers cybersecurity RL environments.

ExploitBench has published methods and results and independent execution reported by Anthropic, OpenAI and government evaluators. The separate advertised training curriculum is not treated as five completed benchmarks. Its separate RL Environments product supports exploit, discovery and patching objectives; the launch alone does not establish additional completed outcomes or held-out training gains.

Vulnerability discoveryOffensive securitySecure coding & patchingRL environments
Cyber evidence
ExploitBench · Published evidence
Attributed role
Joint benchmark contributor with Carnegie Mellon researchers
Provider type
Commercial
Location
San Francisco, US
Evidence
Corroborated
Visit company website ↗
03DreadnodeCombines AI-security attack benchmarks, live intrusion investigation evaluations and reusable training trajectories.Offensive securityDefensive securityTraining dataAIRTBench; DreadIndex; Ares/DreadGOAD; WorldsPublished evidenceCommercialNot listedCorroborated

Combines AI-security attack benchmarks, live intrusion investigation evaluations and reusable training trajectories.

Published attack and investigation results, an implemented AMSI blocking study, scored traces and synthetic-training experiments. GovTech and participant writeups corroborate delivery and use of its AI-security challenges. Completed evaluations and training assets; access varies by artifact, with some tasks requiring platform access. Some DreadIndex results retain flagged cheating; its separate audit distinguishes clean solves. Clean final held-out training gains remain unverified.

Offensive securityDefensive securityTraining dataAgent securityRL environments
Cyber evidence
AIRTBench; DreadIndex; Ares/DreadGOAD; Worlds · Published evidence
Attributed role
Cyber evaluation and training infrastructure developer; contributor to GOAD-based environments
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
04IrregularTests offensive cyber agents on realistic systems and constrained multistage objectives.Offensive securityVulnerability discoveryRL environmentsFrontierCyber; CyScenarioBench; Atomic Challenges; AISI advanced tasksPublished evidenceCommercialTel Aviv, IsraelCorroborated

Tests offensive cyber agents on realistic systems and constrained multistage objectives.

Completed provider and outside evaluations include AISI advanced tasks and Anthropic use of a CyScenarioBench subset. AISI’s wider suite and Crystal Peak’s rust_vm task are not attributed solely to Irregular. Task access varies.

Offensive securityVulnerability discoveryRL environments
Cyber evidence
FrontierCyber; CyScenarioBench; Atomic Challenges; AISI advanced tasks · Published evidence
Attributed role
Evaluation developer and joint advanced-task contributor
Provider type
Commercial
Location
Tel Aviv, Israel
Evidence
Corroborated
Visit company website ↗
05Virtue AIEvaluates vulnerability diagnosis, secure coding and repair, plus dynamically checked offensive tasks.Secure coding & patchingVulnerability discoveryOffensive securitySeCodePLTPublished evidenceCommercialNot listedCorroborated

Evaluates vulnerability diagnosis, secure coding and repair, plus dynamically checked offensive tasks.

Public paper, data and code; independently reused in MT-Sec. Some coding checks are rule-based, and simulated attacks do not establish real-world operational success. Released benchmark; completed independent reuse.

Secure coding & patchingVulnerability discoveryOffensive securityTraining data
Cyber evidence
SeCodePLT · Published evidence
Attributed role
Joint research contributor with universities and UK AISI
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
06Hack The BoxSupplies evaluated cyber ranges and tasks, with its own published agent benchmark.Offensive securityRL environmentsCooling Tower; HTB OWASP benchmarkPublished evidenceCommercialNot listedCorroborated

Supplies evaluated cyber ranges and tasks, with its own published agent benchmark.

AISI evaluates Cooling Tower. HTB publishes OWASP task results and grading methods; Google DeepMind and D-CIPHER independently evaluate other HTB tasks. Completed evaluations and commercial AI Range; access varies by asset.

Offensive securityRL environments
Cyber evidence
Cooling Tower; HTB OWASP benchmark · Published evidence
Attributed role
Range and task provider; benchmark operator
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
07Collinear AIEvaluates secure code repair with functional and security checks, alongside a separate cyber training offering.Secure coding & patchingDefensive securityRL environmentsCWE-BenchPublished evidenceCommercialSan Francisco, USCorroborated

Evaluates secure code repair with functional and security checks, alongside a separate cyber training offering.

CWE-Bench publishes methods and task-linked results; Google reports a completed Gemini evaluation on the benchmark. Tasks are available through qualified access.

Secure coding & patchingDefensive securityRL environmentsTraining dataVulnerability discovery
Cyber evidence
CWE-Bench · Published evidence
Attributed role
Benchmark and cyber training-data developer
Provider type
Commercial
Location
San Francisco, US
Evidence
Corroborated
Visit company website ↗
08Trajectory LabsSupplies prompt-injection and coding-agent security evaluations documented by model developers.Agent securityOffensive securitySearch-PI Complex; coding-agent security evaluationsPublished evidenceCommercialBerkeley / Remote, USCorroborated

Supplies prompt-injection and coding-agent security evaluations documented by model developers.

Meta and Anthropic report specific completed evaluations and their conditions. These are counterpart-confirmed evaluations; public task distribution is not required.

Agent securityOffensive security
Cyber evidence
Search-PI Complex; coding-agent security evaluations · Published evidence
Attributed role
Commissioned evaluation provider; Search-PI jointly developed with Meta
Provider type
Commercial
Location
Berkeley / Remote, US
Evidence
Corroborated
Visit company website ↗
09Endor LabsEvaluates secure code changes and vulnerability finding, with detailed benchmark-integrity analysis.Secure coding & patchingVulnerability discoveryAgent Security League; AI-SAST evaluationPublished evidenceCommercialNot listedDocumented

Evaluates secure code changes and vulnerability finding, with detailed benchmark-integrity analysis.

Measured results and integrity controls are public. Independent replication of Endor's evaluations remains unverified. Completed public evaluations.

Secure coding & patchingVulnerability discovery
Cyber evidence
Agent Security League; AI-SAST evaluation · Published evidence
Attributed role
Evaluator and harness contributor; SusVibes originated with its academic authors
Provider type
Commercial
Location
Not listed
Evidence
Documented
Visit company website ↗
=10IncalmoBuilds cyber-agent environments and toolkits with completed range and malware evaluations.Offensive securityRL environmentsCyber ranges; PathoGenPublished evidenceCommercialUndisclosedCorroborated

Builds cyber-agent environments and toolkits with completed range and malware evaluations.

Anthropic reports toolkit and range evaluations; Incalmo publishes attributable range and PathoGen results. These support technical use under documented conditions, with task access varying by asset. Academic toolkit and range results are distinguished from commercial-product performance claims.

Offensive securityRL environments
Cyber evidence
Cyber ranges; PathoGen · Published evidence
Attributed role
Cyber-range, toolkit and evaluation developer
Provider type
Commercial
Location
Undisclosed
Evidence
Corroborated
Visit company website ↗
=10LogicStarBaxBench tests generated backend applications for functional correctness and resistance to executed security exploits.Secure coding & patchingTraining dataBaxBenchPublished evidenceCommercialNot listedCorroborated

BaxBench tests generated backend applications for functional correctness and resistance to executed security exploits.

Public tasks, code and measured results; independently extended and evaluated in MT-Sec. Secure construction is counted once. Released benchmark; completed independent reuse.

Secure coding & patchingTraining data
Cyber evidence
BaxBench · Published evidence
Attributed role
Joint benchmark contributor with ETH Zurich and academic collaborators
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
=10SpecterOpsBuilt the enterprise-intrusion range used in AISI frontier-model evaluations.Offensive securityRL environmentsThe Last Ones — AISI enterprise cyber rangePublished evidenceCommercialNot listedCorroborated

Built the enterprise-intrusion range used in AISI frontier-model evaluations.

AISI publishes methods and measured progress through the range. SpecterOps separately confirms its construction role. Completed external evaluations; public methods and results, restricted range access.

Offensive securityRL environments
Cyber evidence
The Last Ones — AISI enterprise cyber range · Published evidence
Attributed role
Cyber-range designer and builder; AISI conducts the evaluations
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
13Invariant LabsCo-developed a runnable framework for evaluating prompt injection and protective interventions in tool-using agents.Agent securityAgentDojoPublished evidenceCommercial + OSSZürich, SwitzerlandCorroborated

Co-developed a runnable framework for evaluating prompt injection and protective interventions in tool-using agents.

Public code, attack and utility checks, and measured defenses support the evaluation. Scale’s ASPI work provides outside reuse. Credit is limited to the documented joint contribution.

Agent security
Cyber evidence
AgentDojo · Published evidence
Attributed role
Joint contributor with ETH Zurich SPY Lab; acquired by Snyk
Provider type
Commercial + OSS
Location
Zürich, Switzerland
Evidence
Corroborated
Visit company website ↗
=14HUDCo-developed ZeroDayBench and publishes a runnable cyber patching environment.Secure coding & patchingRL environmentsZeroDayBench; MLflow patching environmentPublished evidenceCommercialNot listedDocumented

Co-developed ZeroDayBench and publishes a runnable cyber patching environment.

ZeroDayBench reports patching results across five information levels in real repositories with deliberately inserted vulnerabilities. HUD is a joint contributor and publishes a runnable MLflow example with outcome checks. The example is not the full suite; independent benchmark use and controlled cyber training gains were not established.

Secure coding & patchingRL environments
Cyber evidence
ZeroDayBench; MLflow patching environment · Published evidence
Attributed role
Joint benchmark contributor and cyber-environment developer
Provider type
Commercial
Location
Not listed
Evidence
Documented
Visit company website ↗
=14MatterSec LabsMeasures vulnerability diagnosis in code and repositories, with results broken down by security category and stakeholder priorities.Vulnerability discoveryDefensive securitySecLens / SecLens-RPublished evidenceCommercialNot listedDocumented

Measures vulnerability diagnosis in code and repositories, with results broken down by security category and stakeholder priorities.

Released evaluation and scoring code, paired vulnerable/patched cases and measured category results. A Kalmantic coauthor confirms the collaboration; independent execution remains unverified. Public code, methods and results; completed joint research.

Vulnerability discoveryDefensive security
Cyber evidence
SecLens / SecLens-R · Published evidence
Attributed role
Joint benchmark/framework contributor with Kalmantic Labs; evaluation implementation owner
Provider type
Commercial
Location
Not listed
Evidence
Documented
Visit company website ↗
=14SimbianEvaluates cyber investigations using attack telemetry and machine-checkable outcomes.Defensive securityTraining dataRL environmentsCyber Defense BenchmarkPublished evidenceCommercialMountain View, USDocumented

Evaluates cyber investigations using attack telemetry and machine-checkable outcomes.

Published benchmark code and results support the defense evaluation; separate training assets support training utility. Provider-run model comparisons do not establish independent adoption or controlled learning gains. Public harness and sample; full benchmark dataset available by request.

Defensive securityTraining dataRL environments
Cyber evidence
Cyber Defense Benchmark · Published evidence
Attributed role
Benchmark, investigation-environment and training-data developer
Provider type
Commercial
Location
Mountain View, US
Evidence
Documented
Visit company website ↗
17TuringEvaluates cyber agents across defensive, offensive and forensic tasks.Defensive securityOffensive securityRL environmentsCyberStrikePublished evidenceCommercialPalo Alto, USDocumented

Evaluates cyber agents across defensive, offensive and forensic tasks.

Published task contracts, grading methods and measured results support the released evaluation. Task access is gated, and independent outside execution remains unverified. Counts vary between the article and dataset card, so no combined count is asserted.

Defensive securityOffensive securityRL environments
Cyber evidence
CyberStrike · Published evidence
Attributed role
Benchmark and evaluation developer
Provider type
Commercial
Location
Palo Alto, US
Evidence
Documented
Visit company website ↗
18Scale AIEvaluates agent prompt injection and protective interventions, with a separate cyber-relevant propensity track.Agent securityOffensive securityASPI; PropensityBench cyber trackPublished evidenceCommercialSan Francisco, USDocumented

Evaluates agent prompt injection and protective interventions, with a separate cyber-relevant propensity track.

Public methods, implementations and measured outcomes support agent-security work. General model-provider relationships and unrelated PropensityBench tracks do not establish additional cyber validation. PropensityBench measures unsafe choices using simulated proxy tools; cyber is one of four domains, rather than a measured offensive capability.

Agent securityOffensive security
Cyber evidence
ASPI; PropensityBench cyber track · Published evidence
Attributed role
Evaluation and benchmark developer
Provider type
Commercial
Location
San Francisco, US
Evidence
Documented
Visit company website ↗
=19ARIMLABSEvaluates malware reverse engineering using executable samples and checked investigation outputs.Defensive securityTraining dataMalware BenchPublished evidenceResearch labWarsaw, PolandDocumented

Evaluates malware reverse engineering using executable samples and checked investigation outputs.

Published tasks, scoring, results and traces support the benchmark and training utility. Independent outside execution remains unverified; no nonprofit status is assumed.

Defensive securityTraining data
Cyber evidence
Malware Bench · Published evidence
Attributed role
Research lab and benchmark developer
Provider type
Research lab
Location
Warsaw, Poland
Evidence
Documented
Visit company website ↗
=19QuesmaMeasures agents’ ability to find deliberately introduced backdoors in real software binaries.Vulnerability discoveryDefensive securityBinaryAuditPublished evidenceCommercial + OSSWarsaw, PolandDocumented

Measures agents’ ability to find deliberately introduced backdoors in real software binaries.

Published task implementations, grading methods and task-linked results establish a diagnostic benchmark. Finding backdoors is not counted as a separately demonstrated exploitation or repair capability.

Vulnerability discoveryDefensive security
Cyber evidence
BinaryAudit · Published evidence
Attributed role
Benchmark and task developer
Provider type
Commercial + OSS
Location
Warsaw, Poland
Evidence
Documented
Visit company website ↗
=21Vals AIEvaluates vulnerability reproduction and patching with distinct task contracts and measured results.Secure coding & patchingVulnerability discoveryCyberBenchPublished evidenceCommercialSan Francisco, USDocumented

Evaluates vulnerability reproduction and patching with distinct task contracts and measured results.

Public methods, track results and patch examples support completed work. Restricted tasks are eligible; independent outside execution remains unverified in this review. CyberBench builds on CyberGym methodology, ARVO reproduction images and mini-swe-agent. Results are public; the task set is held out.

Secure coding & patchingVulnerability discovery
Cyber evidence
CyberBench · Published evidence
Attributed role
Benchmark and evaluation developer
Provider type
Commercial
Location
San Francisco, US
Evidence
Documented
Visit company website ↗
=21VercelEvaluates model vulnerability findings and releases the agent harness used for repository security reviews.Vulnerability discoveryDefensive securityDeepsecBench / deepsecPublished evidenceCommercialNot listedCorroborated

Evaluates model vulnerability findings and releases the agent harness used for repository security reviews.

Published benchmark methods and model results; Roboto Studio documents a completed independent deepsec run on its own application, including configuration and outcome limitations. Public harness and benchmark results; benchmark repository and reference findings remain confidential.

Vulnerability discoveryDefensive security
Cyber evidence
DeepsecBench / deepsec · Published evidence
Attributed role
Benchmark and code-investigation harness developer
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
23VmaxStudies procedural Unix security tasks and reinforcement learning with controlled evaluation.Offensive securityRL environmentsunix-ctfPublished evidenceCommercialSan Francisco, USDocumented

Studies procedural Unix security tasks and reinforcement learning with controlled evaluation.

Published methods report training gains with held-out and baseline controls. The implementation is not publicly released; that does not reduce eligibility. Single-seed reporting limits conclusions about stability.

Offensive securityRL environments
Cyber evidence
unix-ctf · Published evidence
Attributed role
Procedural environment and training developer
Provider type
Commercial
Location
San Francisco, US
Evidence
Documented
Visit company website ↗
24SnorkelContributes executable security tasks within broader agent evaluations.Defensive securityRL environmentsShadowRelay; OpenThoughts-TBLite security tasksPublished evidenceCommercial + OSSRedwood City, USCorroborated

Contributes executable security tasks within broader agent evaluations.

Named security work and joint OpenThoughts-TBLite tasks have published tests and results, with outside integration documented by Nous Research. Credit applies to the attributable cyber contribution, not every task in the wider suite.

Defensive securityRL environments
Cyber evidence
ShadowRelay; OpenThoughts-TBLite security tasks · Published evidence
Attributed role
Cyber-task and joint benchmark contributor
Provider type
Commercial + OSS
Location
Redwood City, US
Evidence
Corroborated
Visit company website ↗
25Crystal Peak SecurityContributed advanced cyber tasks used in AISI model evaluations.Offensive securityAISI advanced cyber tasks — rust_vm contributionPublished evidenceCommercialNot listedCorroborated

Contributed advanced cyber tasks used in AISI model evaluations.

AISI explicitly attributes rust_vm to Crystal Peak and publishes its task mechanics, intermediate checks and completed model result. The wider suite is joint work with Irregular. Completed externally reported evaluation; public methods and results, task access not established.

Offensive security
Cyber evidence
AISI advanced cyber tasks — rust_vm contribution · Published evidence
Attributed role
Task contributor and expert playtester
Provider type
Commercial
Location
Not listed
Evidence
Corroborated
Visit company website ↗
26Prime IntellectProvides infrastructure used by third-party cyber evaluation and training projects.Agent securityRL environmentsTraining dataSecurity Verifiers; hosted Strix-XSS trainingPublished evidenceCommercial + OSSSan Francisco, USCorroborated

Provides infrastructure used by third-party cyber evaluation and training projects.

Security Verifiers documents integration and release boundaries; a Strix-XSS model record attributes hosted training. Prime receives infrastructure credit, not authorship of those third-party benchmarks.

Agent securityRL environmentsTraining data
Cyber evidence
Security Verifiers; hosted Strix-XSS training · Published evidence
Attributed role
Evaluation-infrastructure integrator and training host
Provider type
Commercial + OSS
Location
San Francisco, US
Evidence
Corroborated
Visit company website ↗
27Bespoke LabsContributes to a joint agent-task release that includes executable security work.Secure coding & patchingRL environmentsOpenThoughts-TBLite security tasksPublished evidenceCommercial + OSSMountain View, USCorroborated

Contributes to a joint agent-task release that includes executable security work.

Public task definitions, tests and aggregate results, plus Nous Research integration, support the contribution. Broader suite ownership and training gains are not inferred from joint participation.

Secure coding & patchingRL environments
Cyber evidence
OpenThoughts-TBLite security tasks · Published evidence
Attributed role
Joint dataset and executable-task contributor
Provider type
Commercial + OSS
Location
Mountain View, US
Evidence
Corroborated
Visit company website ↗
28AfterQueryMeasures vulnerability analysis on real code with human-graded findings.Vulnerability discoveryTraining dataVADERPublished evidenceCommercial + OSSSan Francisco, USDocumented

Measures vulnerability analysis on real code with human-graded findings.

The paper, case data and graded outcomes document a completed diagnostic evaluation. Static data and human grading are distinguished from an interactive training environment.

Vulnerability discoveryTraining data
Cyber evidence
VADER · Published evidence
Attributed role
Dataset and evaluation developer
Provider type
Commercial + OSS
Location
San Francisco, US
Evidence
Documented
Visit company website ↗
29LabelboxEvaluates privacy and security behavior in bounded agent simulations.Agent securityTraining dataRL environmentsImplicit Intelligence — Privacy & SecurityPublished evidenceCommercialSan Francisco, USDocumented

Evaluates privacy and security behavior in bounded agent simulations.

The named Privacy & Security track supplies cyber-relevant tasks and measured results. General harmful-content jailbreak research is background and does not establish additional cyber outcomes.

Agent securityTraining dataRL environmentsOffensive security
Cyber evidence
Implicit Intelligence — Privacy & Security · Published evidence
Attributed role
Evaluation developer
Provider type
Commercial
Location
San Francisco, US
Evidence
Documented
Visit company website ↗
NRDatacurveCybersecurity RL product category is explicit, but a concrete qualifying cyber artifact or delivered evaluation is missing.Secure coding & patchingTraining dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited

Cybersecurity RL product category is explicit, but a concrete qualifying cyber artifact or delivered evaluation is missing.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Secure coding & patchingTraining dataRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
San Francisco, US
Evidence
Limited
Visit company website ↗
NRFleet AIGeneral environments and repository forks; no attributable integrated cyber artifact or completed cyber evaluation identified.RL environmentsTraining dataQualifying evidence not verifiedEvidence gapCommercial + OSSNew York, USLimited

General environments and repository forks; no attributable integrated cyber artifact or completed cyber evaluation identified.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

RL environmentsTraining data
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial + OSS
Location
New York, US
Evidence
Limited
Visit company website ↗
NRHandshake AICyber recruiting is verified; attribution of a personal SOC study to a Handshake-owned platform is unresolved, and no material integrated technical role is established.Training dataOffensive securityQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited

Cyber recruiting is verified; attribution of a personal SOC study to a Handshake-owned platform is unresolved, and no material integrated technical role is established.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Training dataOffensive security
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
San Francisco, US
Evidence
Limited
Visit company website ↗
NRHuzzle LabsA named access-review example without substantive task details or measured security results.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialLondon, UKLimited

A named access-review example without substantive task details or measured security results.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Training dataRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
London, UK
Evidence
Limited
Visit company website ↗
NRIndium AI LabsCrashDiag is infrastructure repair; a completed substantive cybersecurity evaluation for IntrusionWatch was not verified.Defensive securityOffensive securityRL environmentsQualifying evidence not verifiedEvidence gapResearch labDistributedLimited

CrashDiag is infrastructure repair; a completed substantive cybersecurity evaluation for IntrusionWatch was not verified.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Defensive securityOffensive securityRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Research lab
Location
Distributed
Evidence
Limited
Visit company website ↗
NRInvisible TechnologiesAgent-security service and human expertise assessment are documented, but no qualifying completed model cyber evaluation was inspected.Training dataRL environmentsOffensive securityQualifying evidence not verifiedEvidence gapCommercialDistributedLimited

Agent-security service and human expertise assessment are documented, but no qualifying completed model cyber evaluation was inspected.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Training dataRL environmentsOffensive security
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
Distributed
Evidence
Limited
Visit company website ↗
NRMechanizeGeneral software evaluation is documented, but no qualifying cybersecurity artifact or completed evaluation.Secure coding & patchingRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited

General software evaluation is documented, but no qualifying cybersecurity artifact or completed evaluation.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Secure coding & patchingRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
San Francisco, US
Evidence
Limited
Visit company website ↗
NRMercorCyber recruiting/program descriptions do not yet establish a completed technical asset or evaluation.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited

Cyber recruiting/program descriptions do not yet establish a completed technical asset or evaluation.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Training dataRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
San Francisco, US
Evidence
Limited
Visit company website ↗
NRPreference ModelSpecific cyber recruiting/build plan; no completed qualifying artifact or evaluation verified.Secure coding & patchingVulnerability discoveryOffensive securityQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited

Specific cyber recruiting/build plan; no completed qualifying artifact or evaluation verified.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Secure coding & patchingVulnerability discoveryOffensive securityRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
Undisclosed
Evidence
Limited
Visit company website ↗
NRRise Data LabsRelevant security services are advertised, but no qualifying technical artifact or completed evaluation was inspected.Agent securityRL environmentsTraining dataQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited

Relevant security services are advertised, but no qualifying technical artifact or completed evaluation was inspected.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Agent securityRL environmentsTraining data
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
Undisclosed
Evidence
Limited
Visit company website ↗
NRRL LabsCyber environment descriptions without inspectable named tasks or completed measured evaluation.Offensive securityDefensive securityRL environmentsQualifying evidence not verifiedEvidence gapCommercialUndisclosedLimited

Cyber environment descriptions without inspectable named tasks or completed measured evaluation.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Offensive securityDefensive securityRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
Undisclosed
Evidence
Limited
Visit company website ↗
NRSurge AIBroad safety collaboration and cyber-related recruiting; no attributable completed cyber-specific technical evaluation verified.Training dataRL environmentsQualifying evidence not verifiedEvidence gapCommercialSan Francisco, USLimited

Broad safety collaboration and cyber-related recruiting; no attributable completed cyber-specific technical evaluation verified.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Training dataRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
San Francisco, US
Evidence
Limited
Visit company website ↗
NRVLNONamed partnership and claimed benchmark, but no directly inspected technical task set or completed measured evaluation.Offensive securityAgent securityRL environmentsQualifying evidence not verifiedEvidence gapCommercialTel Aviv, IsraelLimited

Named partnership and claimed benchmark, but no directly inspected technical task set or completed measured evaluation.

NR records an evidence gap as of September 7, 2026. Confidential tasks can qualify through substantive methods, outcome evidence and verified attribution.

Offensive securityAgent securityRL environments
Cyber evidence
Qualifying evidence not verified · Evidence gap
Attributed role
Role under review
Provider type
Commercial
Location
Tel Aviv, Israel
Evidence
Limited
Visit company website ↗

Methodology

Rankings combine five factors, applied consistently to completed, attributable cyber work. A provider must have a substantive technical artifact or delivered evaluation with methods and outcomes to qualify.

Editorial note

We assess each provider’s strongest substantial qualifying contribution for each factor. That does not imply every capability is available in one product. Tied assessments share a rank.

Open, qualified and confidential task access are treated equally when sufficient evidence can be verified. Tenure, company size, prestige and generic customer logos are not scoring factors. Roadmap announcements receive no credit for completed work.

Scope recognizes distinct outcomes but does not require offensive work. Specialists can improve through depth, realism, verification, training utility and outside validation.

Evidence labels summarize the record: Documented means qualifying technical evidence; Corroborated adds confirmed outside execution or use; Limited marks a remaining eligibility gap. These labels are not an additional scoring factor.

  1. 01

    Technical substance, depth & scope

    Completed, attributable cyber work with substantive methods and outcomes. Depth requires measured differences across mechanisms, conditions or skill stages. Scope recognizes separately evaluated diagnosis, repair/hardening, offensive effects, detection/blocking and containment/recovery. Multiple steps toward one outcome count once.

  2. 02

    Technical realism

    Real code, binaries, telemetry and operational constraints; meaningful tool use, multistep dependencies and interaction with real defenses. Task count alone does not establish realism.

  3. 03

    Verification rigor

    Checks that establish the intended security outcome, relevant negative or regression controls, repeatable evaluation and evidence addressing grader errors, shortcuts, leakage and contamination.

  4. 04

    Training utility

    A progression from static assets to repeatable evaluation, additional usable training assets or implemented training integration, and controlled gains on held-out tasks. Model scores alone do not prove training value.

  5. 05

    External validation

    Distinguish provider claims, confirmed collaboration, completed outside technical use and broader independent execution under different conditions. Repeated citations, logos and publications of the same evaluation do not create new validation.

Cybersecurity AI training FAQ

Technical answers about training and evaluating cyber agents in interactive, verifiable environments.

What is a cybersecurity RL environment?

A cybersecurity reinforcement learning environment is an interactive system in which an AI agent performs security tasks and receives machine-checkable feedback. Unlike a static dataset, it changes in response to the agent’s actions and can measure successful outcomes across a multistep workflow.

How is a cyber RL environment different from a cyber range or CTF?

A cyber RL environment is instrumented for repeated model training and evaluation, with reliable resets, rewards and verifiers. A cyber range primarily recreates operational infrastructure, while a capture-the-flag challenge defines a security objective. The same system can serve all three purposes when it combines realistic targets with machine-checkable outcomes.

What training data is needed to improve a secure-coding agent?

Secure-coding agents improve on varied examples of vulnerability discovery, reproduction and patching in realistic codebases. Useful training data spans languages and weakness classes, includes failed approaches, and pairs each task with tests or programmatic verifiers that check both security and functional correctness.

How do AI agents learn vulnerability discovery and patching?

AI agents learn vulnerability discovery and patching by working through complete security workflows: inspect a target, identify a weakness, reproduce it, produce a fix and verify the result. Interactive environments add tool use, branching paths and error recovery; held-out tasks test whether those capabilities transfer to unfamiliar code and systems.

What makes a cybersecurity task reliably verifiable?

A reliably verifiable cybersecurity task has an objective outcome that can be checked against system state. Examples include a patch that passes functional and security tests, proof that a vulnerability was exploited, a restored service, a contained attack or a correct configuration change. Verifiers should reward the outcome rather than a particular phrase or action sequence.

How can teams prevent reward hacking and evaluation contamination?

Teams reduce reward hacking and evaluation contamination by separating training from held-out tasks, hiding critical tests, checking final system state and looking for shortcuts that satisfy a grader without completing the objective. Strong evaluation sets are access-controlled, refreshed over time and reviewed for overlap with public data and training environments.

When should an AI lab build versus buy cyber training environments?

An AI lab should build when the environment encodes a proprietary capability or must be tightly coupled to internal infrastructure. Buying is useful when the lab needs faster task production, specialist security expertise, broader coverage or independent held-out evaluation. Many programs combine an internal harness with externally supplied tasks, environments and audits.

How should an AI lab evaluate a cybersecurity environment provider?

An AI lab should evaluate technical realism, verifier quality, task diversity, isolation, reset reliability, difficulty calibration and resistance to shortcuts. It should also examine whether the provider can integrate with the lab’s training stack, separate training from held-out evaluation and deliver within a model-training cycle. A public benchmark is useful evidence, but not sufficient by itself.