Gemini 3.8 Live and 3.5 Transcribe land in the API for voice agents

Google's speech-to-speech and transcription models are GA at $0.005/min; plus OpenAI sunsets GPT-5.5 in Codex Oct 14 and a curated market for agent skills.

Nowline SEP 19 12:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Gemini 3.8 Live goes GA for real-time voice

    Google's native speech-to-speech model is generally available on the Gemini API and AI Studio, with an extended-thinking variant that reasons mid-call and asynchronous function calling that fires tools while the agent keeps talking, across 97+ languages. Audio is $0.005/min in and $0.018/min out — roughly $1.38 an hour — so tool-using voice agents are now a config change, not a research project.

  • Gemini 3.5 Transcribe hits 2.6% word-error-rate

    Shipping alongside Live, the new speech-to-text model posts a 4.0% WER streaming and 2.6% batch across 85+ languages, with automatic code-switching and custom vocabulary biasing up to 1,000 domain terms. Accurate enough to drop into call logs, meeting notes, or dictation without a specialist transcription vendor.

  • GPT-5.5 leaves Codex and ChatGPT on Oct 14

    OpenAI is retiring gpt-5.5 from ChatGPT, ChatGPT Work, and Codex on October 14; the OpenAI API is unaffected. If you run Codex with a ChatGPT sign-in, switch to gpt-5.6-sol before then or your agent breaks — worth a config check this week, not on the 14th.

  • Elsewhere: a curated market for agent skills

    Skillbay — pitched as 'Craigslist for agent skills, curated by a human' — hit the Hacker News front page today with vetted, installable skills for coding agents. A weekend way to bolt a battle-tested skill onto your Claude Code or Codex setup instead of writing one from scratch.