The developments most likely to change what AI builders do next.
Models, APIs & Pricing
Google launches Gemini 3.7 Flash for agents
Summary A lower-cost Flash model improves coding, document reasoning, web UI generation, and multi-step tool use.
The details
Gemini 3.7 Flash scored 43.6% on FrontierCode 1.1 Main, versus 34.4% for Gemini 3.6 Flash.
It reached 1,588 Elo on Arena.ai's web-development evaluation, up from 1,538 for 3.6 Flash.
Google reports 34.0% on GDP.pdf versus 22.0%, and 30.4% versus 17.0% on an unnamed business-workflow evaluation.
The introductory API price is $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026; it then rises to $1.50 and $7.50.
Summary A limited-preview OpenAI API tier claims frontier-model output at up to 750 tokens per second.
The details
Cerebras says GPT-5.6 Sol Ultrafast delivers up to 750 output tokens per second in the OpenAI API.
Cerebras reports its Ultrafast configuration completed all 2,500 Humanity's Last Exam questions in 11 hours 11 minutes, versus Claude Fable 5's 78 hours 27 minutes.
On GDP-Val, Cerebras reports a 5.6x end-to-end speedup for GPT-5.6 Sol Ultrafast over standard GPT-5.6 Sol, with no quality degradation in its internal test.
The service is in limited preview for a select group of customers, with access expanding as capacity grows.
Summary Promotional video claims Meta released Muse Code; supplied materials do not substantiate release.
The details
The supplied primary link is a Skool community page, not a Meta announcement, product page, repository, documentation page, or benchmark report.
The video describes a purported command-line agent, Muse Code, with background agents, persistent memory, checkpointing, and Git-worktree isolation.
It claims Muse Code trails Claude Opus in Meta-published comparisons, but supplies neither benchmark names nor scores.
The input combines multiple video descriptions, including unrelated Grokbot-versus-Hermes material, preventing attribution or release details from being resolved from the source.
Summary Blomfield urges companies to automate feedback, evaluation, and code changes around measurable outcomes.
The details
Blomfield describes a YC data-query agent and an overnight agent that reviews queries and opens pull requests for recurring failures.
He defines an AI loop as real-world inputs, a policy layer, tool access, quality gates, and learning tied to measurable outcomes.
The talk recommends automated or model-based quality gates for routine work, reserving human approval for exceptional cases.
Blomfield expects AI could potentially handle YC’s application-to-advice workflow end to end by late 2026 or early 2027; deployment remains unconfirmed.
Summary Walkthrough proposes persistent style files to reduce jargon and control response length.
The details
The walkthrough says Claude Code offers default, proactive, explanatory, and learning styles, plus custom styles configured through its settings UI.
It proposes a custom ELI5 output style and ASD-STE100 constraints to reduce jargon and use restricted technical vocabulary.
The suggested setup creates a style file, changes relevant settings, then starts a new session to confirm the style is active.
Evidence limit: supplied material includes no Anthropic announcement, documentation link, exact file path, token-cost measurement, or controlled evaluation of claimed behavior.
Summary Demo pairs a local-file browser UI with a small server.
The details
The claimed implementation uses an HTML page, a small server, and conversations stored as plain files in a local folder.
The first version polls for replies on a 60-second loop.
Later, the creator describes persistent sessions with unique session IDs and a watcher.
The demo compares an empty folder with an established myPKA folder containing CLAUDE.md and referenced documents for instructions and accumulated context.
The creator says the setup uses a Claude subscription, not a per-token API, but provides no cost comparison or source code.
Summary A shared agent skill generates self-contained editorial HTML and SVG diagrams, with static output as default.
The details
The 27 diagram types have minimal-light, minimal-dark, and full-editorial static HTML variants requiring no build step, JavaScript, or external images.
Version 2.3 adds semantic system patterns and optional accessible motion, while static output remains the default.
The skill redraws draw.io or Mermaid inputs at a selected format, size, and detail level.
It installs via Claude Code, Codex, and Pi marketplaces; onboarding maps a site's palette and fonts to semantic tokens, proposing a style-guide diff.