Agents & Gemini API, MTS @GoogleDeepMind | prev: Tech Lead at @huggingface, AWS ML Hero 🤗 Sharing my own views and AI News 🧑🏻‍💻 philschmid.de

Nürnberg
The Interactions API is now generally available. 🎉 The Interactions API is the simplest way to build with Gemini for humans and agents. - One API for Gemini models and agents. - Antigravity Agent with a isolated remote Linux sandbox. - Image gen with Nano Banana; music with Lyria 3, soon video with Omni. - `background=True` for async, long-running interactions. - Multimodal Tool Use & Combination. - Dedicated skills for coding agent. Building something new inside Google isn't easy, but it's possible. We spent almost a year on this because we think developers and agents deserve an API that is intuitive, familiar and easy to learn and use and evolves with improving capabilities and new behaviors. Give it a try. Feature requests, bugs, complaints all to me. 👇🏻
27
23
266
36,153
Powered by Gemini 3.8 TTS.
To our most astute listeners, if you've noticed a fresh change in our sound— you're spot on. Our AI hosts would love to explain what's happening (and what you can look forward to)!
5
6
149
18,394
With Gemini 3.8 TTS you can replicate your voice or design a completely custom one from a prompt 1. Record 20s of you talking + the consent sentence 2. Create your Voice via API call 3. Use it in any request, style goes in speech_metadata Past this into your agent "Read philschmid.de/gemini-3-8-tts and walk me through creating my own voice for Gemini 3.8 TTS. Check my setup first (GEMINI_API_KEY, ffmpeg, gemini-skills), help me record the two clips, create the voice, and generate a test line I can listen to." or read below.
22
32
326
25,459
Gemini 3.8 Flash. Just use Gemini for multimodal understanding.
I drew this today. None of the frontier models come anywhere close to matching the correct name to each person. I feel like this is a pretty good visual test so I’m looking forward to trying it with future models.
13
2
204
20,006
Gemini 3.8 Flash TTS and Flash-Lite TTS are here.🗣️ You can now replicate your voice, design new ones from a prompt and direct every line. Gemini ranks #1 on Hume's Voice Design Benchmark and tops Voice Arena in 6 languages. 🎙️ Replicate a voice from 30s of audio, or describe one in a sentence 📚 Voice library filterable by language, accent, gender, pitch, persona 🎭 Direct line by line: style, <laugh>, <sigh>, |backchannels|, 2-speaker scenes 🏆 #1 Hume Voice Design, #1 + #2 Hume Quality, top of Voice Arena in 6 languages 🔒 Consent check on replication, SynthID on every clip Create your voice once and use it in any TTS call. Available in the Gemini API and @GoogleAIStudio.
19
15
173
15,153
Love all these Jev demos, replications and projects. They show how much becomes possible with ultra-fast, multimodal models that are cheap enough to use freely. We are so focused on coding agents that it’s easy to forget how much else AI can do.
31
14
235
11,653
Excited to share that we partnered with @speakeasydev to build the sdks for the Interactions API and supported them to open-source their their entire OpenAPI generator suite (SDKs, agent CLIs, and MCP servers) open source for the community to use: developers.googleblog.com/wh…
7
8
55
7,192
Today we are releasing a new Gemini managed agents version with better caching, lower costs, a Files API, and a Credentials API. 📉 Up to 30% lower costs with up to 22% higher cache hits 📁 Files API: Upload, list and download files from sandbox easily 🔐 Credentials API for MCP servers, OAuth2, and 3rd-party APIs as secure egress proxy 🆓 Free tier to experiment in @GoogleAIStudio & Gemini API Start Here: aistudio.google.com/learn/ma…
23
23
321
57,953
New Credentials API to securely authenticate with external services and MCP servers without the model ever seeing your secrets: - Register secrets once with `credentials.create()` supporting `bearer_token`, `environment_variable`, or `oauth2` (with auto-refresh and token rotation). - Zero plaintext exposure: Secrets never enter model context, stdout, memory. - The agent only sees a placeholder (e.g., `__GEMINI_CRED_slack-bot-token__`). - Egress proxy swaps real token on the request only for allowlisted `trusted_domains` (any exfiltration attempt to another host gets a `403`). - Bind credentials directly to remote `mcp_server` tools or sandbox env vars. Docs: ai.google.dev/gemini-api/doc…
4
1
8
1,875
New Files API to move data in and out of your agent environment: - Seed files inline or upload data mid-conversation with `files.upload()`. - Inspect generated files and directory sizes with `files.list()`. - Download artifacts, dashboards, reports, CSVs, or git repos with `files download()`. - Sandboxes persist across turns via `environment_id` Docs: ai.google.dev/gemini-api/doc…
2
3
1,426
We just launched the DeepMind Institute, a platform for researchers across @GoogleDeepMind and the wider research community to publish and debate how increasingly capable AI should be built, governed and used. First 4 essays on reasoning transparency, economic policy and human flourishing at institute.deepmind.com.
16
26
273
11,056
Gemini 3.8 Live and 3.8 Live Extended Thinking are here. Following 3.5 Transcribe last month, this continues our focus on real-time voice agents. 🐸🐸 - 82.6 (#1) on Artificial Analysis Quality Index - $0.005/min input and $0.018/min output. - 35.1 (#1) Agentic task completion (τ-banking) 3.8 Live Extended Thinking can think in the background and run asynchronously tool calls while keep talking to you. It uses background thinking to coordinate multi-step tool use without losing conversational flow. Available in @GoogleAIStudio Gemini API, Google Search, and @GeminiApp or as partner plugins in @LiveKit, @Pipecat_ai, @LangChainAI, and @vercel
11
11
174
10,761
output_schema + atateless + code mode is going to drive a MCP comeback.
8
1
33
5,163
What a weird weekend. Good thing tomorrow is Monday.
20
6
372
25,942