How Model Evolution Has Reshaped Security AI Agents on H1 Platform
As models have rapidly crossed capability thresholds, the harnesses built around AI agents on H1 Platform have transformed. Each time a newer, smarter LLM model has been released, we've found ourselves removing layers of code that existed to compensate for the previous generation's limitations. This article describes how our agent harnesses have drifted toward architectural simplicity and how our approach to guardrails changed with each generation.
What changed in the last 18 months
We went through the iteration history of a subset of AI agents that HackerOne runs for code analysis and security testing. These agents are designed to review pull requests for vulnerabilities, validate researcher-submitted vulnerability reports, analyze whether vulnerable dependencies are actually reachable in customer code, and audit codebase estates for exploitable flaws. They operate on customer code within the scope customers configure, at production scale, with findings posted to pull requests, to code security audit programs, or used to conduct vulnerability root-cause analysis.
We tracked different dimensions across the repositories used to maintain the agents: model selection (which model version, and why it changed), output format (how the agent returns structured findings), control flow (how the harness constrains what the model does), and guardrails (what reduces the risk of bad outputs from reaching customers).
The timespan this document covers marks what we're calling our ‘modern AI harness era’: March 2025 through August 2026. This describes when we began actively winding down models we'd built in-house like Security Hotspots, built off of CodeBERT and trained with participation from a group of PullRequest open source customers, and converted our monolithic AI pipelines to hierarchical orchestrations where multiple agents carry out full codebase surface navigation in sequential workflows. Today the AI agents we design and build are based on core functions of real-world continuous threat exposure management (CTEM) and evaluated based on how proficient and consistent they are at doing their job.
HackerOne does not use or permit use of confidential researcher submissions or customer vulnerability data to train, fine-tune, or otherwise improve generative AI models.
Models didn't move in one direction
The fleet didn't uniformly upgrade. Different agents moved in different directions based on their task shape.
In our testing, one of the validation workflow agents for code security analysis saw a performance gain moving away from higher-capability, higher-latency models like Opus to balanced, faster models like Sonnet. What happened was that the output from judgment-focused agents it worked off of, its “teammates”, became more consistently accurate while, in parallel, new generations of the balanced-class models were advancing. This meant that the agent could complete its task to arrive at an actionable verdict decision faster without sacrificing accuracy. In our internal testing, heavier high-capability models were comparably accurate on this task but added a noticeable latency. For our H1 Code product this was especially important; every extra second a CI/CD pipeline check runs degrades the user experience.
Our orchestration agent, which coordinates multiple analysis phases and renders a final judgment, moved up to deeper reasoning models. Raising the reasoning ceiling substantially improved the effectiveness of discovery agents working downstream.
Several multi-agent workflows that started exclusively on higher-reasoning models for all phases shifted to faster, balanced models for their breadth work (scanning, exploration, information gathering) and kept the deeper reasoning models for judgment-heavy tasks like verdicts and customer escalation paths. For example, in our testing, the entire code auditing workflow saw a performance gain when a batch scanning agent, used for reconnaissance activities where throughput and speed matter more than reasoning depth, was assigned to faster balanced models. Then saw incremental improvements in our benchmarks correlating with the evolution of Sonnet from 3.5 to 3.7 to 4.5.
By the end of August 2026 most ‘breadth tasks’ (e.g., scanning, routing, exploration) performed best on balanced, faster models in our benchmarks. ‘Judgment tasks’ (verification, final verdict, orchestration) performed better with deeper-reasoning models.
Working with frontier cyber AI models
The importance of model-to-task delegation and empirical proof we’d started to see was amplified as we began working with a new generation of frontier cyber AI models like Mythos 5 through Anthropic’s Project Glasswing and GPT cyber models through OpenAI’s Daybreak Defense Network. These models introduced new concepts like compositional risk: vulnerabilities that don't live in any single commit, but in how multiple safe-in-isolation changes interact with the code base to create complex vulnerabilities over time.
Our work with these models has led to new dimensions of task proficiency in our secure code analysis harnesses and the benchmark sets we evaluate them on.
For example:
- Discovery involving deep investigation of an adversary attack path thesis.
- Remediation strategy that is clear, actionable, and acknowledges context-relevant constraints.
Agents followed a similar output-format arc
This played out across three phases, separated by about six months each. Every agent went through all three, just on different timelines.
Phase 1: Free-form JSON (2025). Ask the model to respond in JSON. Parse (hopefully). Handle malformed responses with retries. Our oldest agent ran this way at first. During that period, a measurable fraction of runs failed silently because the model returned valid JSON that didn't match the expected schema. Unfortunately we were catching these in production logs, not at the API layer.
Phase 2: Forced structured output (late 2025). API-level structured output (forced tool calls or JSON schema mode) moved validation from heavy-handed prompting in code to the model. Malformed responses became loud API errors, not quiet data corruption errors.
Phase 3: Tool-call-as-output-channel (2026). Our most mature agents use tool calls not for actions but as typed output ports. The agent defines schema-validated tools whose sole purpose is to emit structured findings. This is called an output contract. The model explores freely within containment (reads files, greps for patterns, navigates the codebase), then emits findings through those typed channels. This decouples exploration from reporting: the agent doesn't need to hold its full analysis in a single response.
Each phase was enabled by the model below, becoming reliable enough to graduate. When models reliably produced the JSON format we needed, we moved validation to the API. When they reliably called tools, we could give them a tool schema to use as a contract between them.
Invocation
Our agent fleet cycled through four different software methods for calling AI models over the 18 months:
- Direct API calls (early 2025). Raw model invocation.
- CLI subprocess (mid-2025). This didn't last long before we replaced it with a proper SDK.
- Agent SDK (mid-to-late 2025). Four of our core agents adopted it. One removed it after just a few weeks; the others migrated away over the next six months.
- Graph framework + direct model calls (2026 onward). The current convergence: a graph execution framework owns the workflow structure, direct API calls own model invocation, and middleware owns the guardrails.
The agent SDK was rolled out for four of our core product agents. It was replaced because no single SDK spanned all three concerns the team needed to control independently: graph structure, model invocation, and enforcement middleware.
In 2026 we added an invocation layer that routes to multiple providers at runtime. Some model families support thinking blocks and prompt caching; others don't. A factory layer abstracts the difference so the harness code above it doesn't know or care which model answered.
Guardrails got stricter as models got smarter
With each model generation, we saw fewer formatting errors and hallucinations and entrusted agents to "cook" more. This opened up new categories of failures that required new guardrails.
2025: Implicit trust. Sequential processing, no turn limits, retry-until-success. The model ran until it finished (or crashed).
Early 2026: Soft budgets. Fixed thinking-token caps. Turn limits as backstops. These existed to catch runaways, but not to shape behavior.
Mid-2026: Hard budgets and stuck-turn detection. Agents added two-tier tool-call limits: a soft limit that injects a convergence message ("wrap up your analysis"), and a hard limit that terminates the run. Some agents added per-phase budgets, giving exploration phases fewer tool calls than deep-analysis phases. Middleware began detecting when the model was repeating itself and injecting explicit pressure to converge.
August 2026: Failure policy split by task. Our orchestration agents abort an entire run on an incomplete model turn. Other agents salvage partial work on errors rather than discarding it. Truncated and empty model turns that provide no value are rejected.
Capable models are more creative but can roam without making progress. If unguided, an agent will find local optima of "look busy" we didn't see from earlier models. The guardrails prevent agents from finding an efficient way to stall.
Better models, fewer architectural layers
In multiple instances, each time a model crossed a capability threshold, we pruned architectural layers as if they were technical debt.
Agent specialization became less important. One of our agents originally fanned out analysis work to category-specific specialists (e.g., injection flaws, auth flaws, infrastructure misconfig, and so on). On our internal test set, a single generalist agent with a property-based prompt began to outperform the specialist fan-out. We ultimately removed the specialist layer, hundreds of lines of orchestration and routing code. The specialist layer existed because agents using the earlier models couldn't hold all security dimensions in a single pass.
Scaffolding fields became less necessary. One agent originally required the model to output reasoning summaries alongside its scores, forcing it to "show its work" as a way of improving output quality. In our evaluations, we stripped those fields and it produced better scores.
This is different from coverage visibility. The agents still record what code they examined and what they skipped. The scaffolding that impacted accuracy was the forced narration.
Extended thinking became less necessary (for some tasks). One agent disabled extended thinking for batch jobs after testing showed it was consuming wall-clock time without improving accuracy. Smaller thinking budgets or none at all produced the quality within a fraction of the time.
In short, almost every major model generation absorbed one layer of compensating complexity. Specialists became generalists. Multi-turn became single-call. Scaffolded reasoning became structured output. The agents got simpler.
Building AI agents for security
Model changes over generations force architectural change. Every decision about multi-agent composition, output scaffolding, and SDK choice has a shelf life measured in model generations. Avoid building for permanence.
The two-tier model split (a smaller model for breadth, a reasoning model for judgment) has held stable for six months, which, in our experience, is a realistic ceiling for how long a harness design survives before it needs revisiting. If an agent does both scanning and verification, separate concerns, assign different models to each, and measure results.
The need for strict guardrails doesn’t wane with new model generations. The failure modes of higher-capability models are subtler and harder to detect with each generation than their predecessors.
See it in action
The agents described span three products on the H1 Platform, each aligned to security programs and continuous threat exposure management (CTEM) workflows organizations maintain.
- H1 Code runs pull-request and CI/CD checks designed to catch new risks before they merge.
- H1 Remediation turns confirmed findings into fix strategy.
- H1 Code Security Audit runs a broad-coverage audit using a hierarchical orchestration of scanning, exploration, and verification agents that work from reconnaissance to verdict.
Want to see what frontier cyber AI models with an architecture built for them can surface in your attack surface?
Check out H1 Code Security Audit