Source-led article

Sakana AI Launches Fugu-Cyber Model With Specialized Security Reasoning Endpoints

AI News India//3 min read
A representation of AI security orchestration and code analysis on a dark dashboard display.
A representation of AI security orchestration and code analysis on a dark dashboard display.
Journalists Protest against rising violence during march in Mexi | by Knight Foundation | openverse | by-sa

Sakana AI has expanded its Fugu orchestration architecture with the launch of Fugu-Cyber, model ID fugu-cyber-v1.0, a dedicated endpoint focused on cybersecurity reasoning and agentic validation. The release follows the initial deployment of the Fugu orchestrator and positions the new model alongside competing enterprise cyber security reasoning frameworks. Sakana reports benchmark scores of 86.9% on CyberGym and 72.1% on CTI-REALM, placing its performance alongside specialized frontier offerings in vulnerability analysis and threat intelligence mapping.

Evaluation Context and Benchmarks

The reported success rate of 86.9% on CyberGym positions Fugu-Cyber near recent frontier benchmarks in automated vulnerability assessment. When initial CyberGym evaluations were published, standard agent-model pairings achieved roughly 20% success rates. Subsequent evaluations by Anthropic for Claude Mythos Preview under Project Glasswing reached 83.1% in April 2026, while OpenAI reported 85.6% for its updated GPT-5.5-Cyber. Sakana’s 86.9% score represents a modest advance on the established frontier rather than a wide architectural leap.

The CTI-REALM evaluation presents a different metric structure, scored as a trajectory reward between 0 and 1 rather than a binary pass or fail rate. Microsoft internal evaluations previously placed top configurations, all utilizing Claude, in a band from 0.624 to 0.685. Fugu-Cyber records a 72.1% rating in this framework, though direct comparisons require caution due to differences in trajectory scoring across testing environments.

Orchestration and Verification Architecture

Fugu operates as a language model trained to parse a given user query and dynamically build an agentic execution scaffold, delegating sub-tasks to specialized models within a managed pool. This workflow draws on documentation from the Fugu technical report and ICLR 2026 research papers, including TRINITY and the Conductor. TRINITY assigns specific Thinker, Worker, and Verifier roles across multiple large language models, while the Conductor utilizes reinforcement learning to optimize coordination strategies in natural language.

For cybersecurity applications, the verification layer acts as a primary control point. Candidate vulnerabilities identified by initial scanning agents undergo secondary validation by security-specialized sub-agents before any patch or remediation step is proposed. Because internal routing pathways remain proprietary, developers cannot inspect which specific underlying model handles each intermediate step of the execution trace.

Datos clave

Metric | Value | Context
CyberGym Success Rate | 86.9% | Comparable to reported frontier models
CTI-REALM Score | 72.1% | Trajectory reward score across threat workflows
Base API Pricing | $6 / $36 per million tokens | Rates double above 272K context window

Access Controls and Pricing Structure

Deployment of Fugu-Cyber is subject to a strict multi-dimensional gating process. Developers must submit an application form detailing their intended use case and verified contact information, with each request undergoing manual review by the Sakana team. The model operates under an updated Acceptable Usage Policy that expressly prohibits offensive cyber misuse.

Billing is restricted to the Token Plan. General subscriptions such as the $20, $100, and $200 tiers cover standard Fugu and Fugu-Ultra endpoints only. Pricing for Fugu-Cyber is fixed at $6 per million input tokens, $36 per million output tokens, and $0.60 per million cached input tokens. Furthermore, all three rates double once a request crosses a 272K-token context window, making the higher tier relevant for extended codebase analysis. Additionally, the Fugu API remains unavailable in the EU and EEA while the company aligns operations with GDPR compliance requirements.

Source: MarkTechPost, https://www.marktechpost.com/2026/07/25/sakana-ai-releases-fugu-cyber-orchestration-model-cybergym-cti-realm/

Datos clave

Punto Detalle
Fuente MarkTechPost
Fecha 2026-07-26T00:12:57+00:00
Tema Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM