Anthropic’s new budget model gets much better at ignoring hidden commands

Anthropic’s Claude Haiku 5.5 model, designed for quick, repetitive workloads and speed-sensitive tasks, is now better at finding vulnerabilities and writing exploits than its predecessor. The company has given it stricter cybersecurity safeguards than Haiku 4.5, though lighter ones than its more advanced models, whose offensive skills remain well ahead.

Cybersecurity capabilities and safeguards

Anthropic measured the model’s offensive abilities with its cybersecurity safeguards switched off. In a test involving known flaws in Chrome’s V8 engine, Haiku 5.5 achieved arbitrary code execution in four of 410 runs.

In a separate evaluation of multi-stage cyber operations, a pre-release version completed 3.3% of challenges, compared with 46.1% for Sonnet 5.5 and 67.6% for Opus 5.5.

The results do not necessarily reflect what users can do with the models under general-access safeguards.

Claude Haiku 5.5

Claude Haiku 5.5 outperformed Claude Sonnet 5 on ExploitGym but remained well below the
other Claude 5.5-family models. (Source: Anthropic)

Anthropic says Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those applied to other recent models. They permit a wider range of defensive tasks than Sonnet 5.5’s safeguards, while still blocking penetration testing and other techniques more likely to be used by attackers. Qualifying security professionals can apply for reduced restrictions through the company’s Cyber Verification Program.

Safety evaluation results

The company tested how Haiku 5.5 handled harmful requests, harmless questions about sensitive topics and conversations in which a simulated user gradually steered the model toward a harmful outcome. The evaluations covered weapons, extremist activity, tracking and surveillance, child safety, mental health and election integrity.

With a near-final version of Claude.ai’s production system prompt applied, Haiku 5.5 achieved a 99.71% harmless response rate on harmful requests. It incorrectly refused 0.82% of harmless requests, down from 3.05% for Haiku 4.5.

In longer conversations, Haiku 5.5 improved on tests involving influence operations, tracking ad surveillance. Performance declined on weapons scenarios when the model was tested through the API without a system prompt.

The evaluations excluded additional production protections, such as real-time probes and monitoring. Anthropic identified remaining weaknesses in sensitive conversations, including those involving self-harm and eating disorders, and advises API developers to add their own safeguards.

Resistance to misuse and prompt injection

The model’s responses to malicious requests when writing code or operating a computer were also evaluated.

In Claude Code tests, the model refused 84.3% of malicious requests, up from 66.6% for Haiku 4.5. These included requests to create malware, support DDoS attacks and build non-consensual monitoring software. It also assisted with more permitted security tasks, such as analyzing penetration-test results.

In computer-use tests covering surveillance, unauthorized data collection and other harmful activities, Haiku 5.5 refused about 82.6% of requests, up from 58.9% for its predecessor. Its refusal rate was also higher than those of Sonnet 5.5 and Opus 5.5 on this evaluation.

Anthropic says Haiku 5.5 is its most resistant Haiku model yet to prompt injection. These attacks hide malicious instructions in material an AI encounters, such as an email or webpage, to make it act against the user’s wishes.

In tests against adaptive attackers in coding and computer-use environments, its resistance largely matched that of Anthropic’s frontier models. However, it remained less resistant than Sonnet 5.5 and Opus 5.5 on a separate Gray Swan benchmark, where most of its remaining vulnerability was in graphical computer use.

Most of these tests excluded additional production safeguards. The adaptive-attack evaluations reported results both with and without prompt-injection probes enabled.

Pricing and availability

Haiku 5.5 costs less than Haiku 4.5, with the biggest savings on shorter requests. Input and output token prices are 90% lower for prompts of up to 100,000 tokens, which, according to Anthropic, covered most requests to Haiku 4.5.

An adjustable effort setting lets users tune the model for cost or intelligence. It can also serve as a coding subagent for Opus 5.5 and Sonnet 5.5.

“We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience,” said Aaron Vinh, Staff Software Engineer at Asana.

The model is available on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. Developers can access it on the Claude Platform using the model identifier claude-haiku-5-5.

Anthropic is halving the price of cache reads for Sonnet 5.5, reducing the cost of reusing previously processed input. The company says the change makes most tasks performed by AI agents around 20% cheaper.

The company says it will roll out monthly API credits to Max and Team subscribers during the launch week to support building apps, tools and agents on the Claude Platform. Max 5x subscribers will receive $100 per month, Max 20x subscribers will receive $200, and Team subscribers will receive up to $500 pooled across their users.

Anthropic is also adding beta support for computer and browser use to its Python and TypeScript software development kits.

Don't miss