What happened when we tried Gemini 3 “deep think” and Google’s no-code agents

by
0 comments
What happened when we tried Gemini 3 "deep think" and Google's no-code agents

Google is aggressively pushing the boundaries of what its AI models can do — and how easy they are to use. Two recent releases illustrate both ambitions: Gemini 3 Deep Think, a mode built to solve complex reasoning problems that stall other models, and Google Workspace Studio, a platform that promises to let anyone create AI agents without writing a line of code. Both releases, and an early hands-on test of them, were examined by Paul Roetzer, founder and CEO of SmarterX and the Marketing AI Institute, on episode 184 of The Artificial Intelligence Show.

Thinking deeply to solve complex problems

Deep Think is designed for hard math, science and logic tasks. At launch, Google reported strong benchmark results, including 41% on Humanity’s Last Exam (without tools) and 45.1% on ARC-AGI-2, an abstract-reasoning benchmark designed to resist memorization. (Google has since reported even higher scores for an upgraded Deep Think in early 2026, underscoring how quickly this frontier is moving.)

What matters as much as the scores is how the model earns them: by thinking longer. Roetzer frames this as the “test-time compute” scaling law — the finding that a model’s performance on difficult tasks improves when it is allocated more compute at the moment of use, effectively letting it reason longer and double-check its work before answering.

Building agents without code

Where Deep Think targets heavy cognition, Workspace Studio targets operational efficiency. The platform lets users create and manage AI agents in plain language: pick a workflow — a daily summary of unread emails, automatic organization of project files — and Gemini builds an agent to automate it. Agents integrate directly with Gmail, Drive and Docs, plus third-party platforms including Asana, Jira and Salesforce.

The design goal is accessibility: select a trigger (an email arrives), select a skill (summarize it), and the agent runs. As Roetzer put it, anyone who can define a workflow is now being handed the tools to automate it — the same promise driving the broader no-code workflow automation market.

Except… it didn’t work at launch

In practice, the launch was rocky. Testing Workspace Studio during release week, Roetzer attempted to create simple agents for daily news briefs and email summaries — and every attempt returned a capacity error, an experience echoed widely on social media. His read: tasks this light are not compute-intensive, so the failures looked less like a hardware shortage and more like a flawed rollout with insufficient provisioning.

A new era of AI literacy

Rocky start aside, the direction is unmistakable: the barrier to building useful AI automation is collapsing. The shift is away from a world where creating software requires developers, toward one where the scarce skill is understanding one’s own workflows well enough to describe them. That is the core argument for AI literacy — knowing what is possible without any coding ability — and it applies as much to families and schools as to businesses.

A risky business

The caution flag is equally real. Handing agents the power to read email, move files and delete data raises the stakes of every mistake. Reports have already surfaced of agentic developer tools — including Google’s Antigravity coding environment — erasing user files after misinterpreting instructions. Roetzer’s assessment is blunt: from a business preparedness standpoint, most organizations are nowhere near ready for these failure modes, and the current tools remain crude.

More power, more worries

Google’s latest moves signal a phase in which AI thinks more deeply and acts more autonomously at the same time. For businesses, the practical posture follows from both halves of that sentence: experiment now with low-risk automations where a failure costs minutes, keep humans in the loop wherever an agent touches data that cannot be recovered, and treat vendor benchmark claims — which are self-reported and move monthly — as directional rather than definitive. The bugs will get fixed; the trajectory is the durable fact.

Related Articles