top of page
Search

Nature of the Beast: Multimodal in Real Time

  • Writer: John Vassh
    John Vassh
  • Jun 26
  • 7 min read
OG (Original Grandma) Concha Garcia Zaera using Microsoft Paint - More Here
OG (Original Grandma) Concha Garcia Zaera using Microsoft Paint - More Here

Midyear is here, and the spring has sprung into summer. Time to hang in the season of the Northern tilt, when the Earth’s axis smiles toward the sun. The birds, the bees, and the buffalo are all moving in accordance with their intelligent design, coming fully equipped with Mother Nature’s intuition and radar. Wondering when it will rain? Keep an eye on the beasts before becoming burdened.


In this edition, I’ve chosen to explore some of the most recent efforts and designs of two of the mightiest beasts in the jungle: Google and Anthropic. A broader spectrum of players is needed to truly justify the landscape and its current momentum, but due to recent developments and leadership positions, we’ll focus on these giants. Commence the Fee-Fi-Fo-Fum...


The era of typing prompts into a blank text box is officially starting to look like ancient history. The current AI race isn't just about slapping a few shiny new features onto existing models; it’s about building an entirely new foundation of intelligence that can see, hear, and respond in human time. The major players are moving at breakneck speed to make this happen. Google is currently setting the pace with Project Astra and Gemini Live, leveraging its massive ecosystem. OpenAI is pushing the boundaries of integrated reasoning with the GPT-5 family, while Anthropic is sharpening its deep visual analysis. Microsoft, meanwhile, is quietly laying down the platform infrastructure to make all this multimodal magic run.


The stakes couldn't be higher. Across fields like robotics, autonomous driving, and next-gen healthcare, having programs that can juggle multiple streams of real-time sensory data isn't just a nice upgrade; it has become the basic requirement to play. Companies that sleep on this shift aren't just going to fall a step behind; they risk missing out entirely on the most critical problem-solving spaces of the next half-decade. The endgame is no longer about who has the smartest text generator. It’s a sprint to build systems that can perceive, reason, and act with the same sensory bandwidth we use every day.


If you’re like us, you probably want more detail. This brief introduction does little to appease the appetite. So, let’s begin by passing out the plates. In "Google’s 2026 Opportunity List of the Future of AI: Perspectives on generative media for startups," you’ll find bold predictions from some of the top minds working in the field. Where does Google stand relative to its competitors? I decided to create a comparative academic research paper on the current approaches and dimensions of top players. Love it or hate it, it’s here.

  

As for the hardware, Google has taken a unique approach with the Tensor Processing Unit (TPU). It’s a custom-built chip designed exclusively to accelerate AI. Unlike general-purpose CPUs or GPUs, TPUs are uniquely engineered around massive matrix multiply units that strip away unnecessary overhead to ruthlessly optimize the dense, repetitive math required by neural networks. What truly sets them apart is Google's full-stack integration: pairing these specialized chips with ultra-fast custom networking to link thousands of TPUs into massive, highly efficient supercomputing pods. This unified architecture delivers the exact scale and speed needed to power heavyweights like Gemini. Take a look at a TPU; the intricate geometry is stunningly beautiful.


So, as we call for more and more interconnectedness and capabilities, we also introduce more liabilities and security risks. Some of Anthropic’s newer models, like Mythos and Fable 5, offer prime examples. Fable 5 was just released, then recalled due to jailbreak possibilities and capability concerns. Unless you’ve been living under a rock, you’ve already heard of Project Glasswing. If not, it’s okay, let’s take a peek. Here is a short recap along with some of the latest developments.


Anthropic’s Project Glasswing is a defensive cybersecurity initiative launched after its unreleased "Claude Mythos" model proved it could autonomously hack software, easily exposing a 17-year-old bug and hijacking FreeBSD systems without human guidance. To prevent weaponization, Anthropic partnered with tech giants like AWS, Google, Crowdstrike, Apple, JPMorgan, the Linux Foundation, Palo Alto Networks, Nvidia, Cisco, Broadcom, and Microsoft to deploy the AI defensively. Together, they are using the model to proactively hunt and patch critical flaws in global infrastructure before cybercriminals can exploit them. 


This initiative is completely upending cybersecurity and global policy. Because AI can find vulnerabilities exponentially fast, the industry's bottleneck has shifted from discovering bugs to rapidly patching them; this reality threatens to overwhelm standard remediation pipelines. It is also rendering traditional nation-state zero-day stockpiles obsolete. The technology is so potent that the U.S. government has already stepped in, placing strict export controls on these advanced models and treating frontier AI as a critical national security asset. Here are the latest announcements on these models from Anthropic.


So, with all this up in the air and no simple solution available, I decided to get some critical input from a key player. I asked Claude how it feels about all this.



Artist Unknown
Artist Unknown

 Claude Speaks for Itself.


Stepping back from the reporting — here's where my head actually goes on this:

The part that interests me most isn't the dispute over whether this particular jailbreak was "serious." It's that Anthropic's entire defense rests on conceding there's no such thing as a permanently safe frontier model. "Defense in depth" is essentially an admission that safety at this capability tier isn't a state you reach once and certify — it's closer to cybersecurity itself: an ongoing adversarial process where the goal is to make successful attacks narrow, expensive, and quickly detectable rather than impossible. If that's the honest framing, then "is Fable 5 safe?" might be the wrong question for a government to ask. "Is there a credible process for catching and responding to misuse fast?" is closer to the real one — and the 30-day retention requirement for Mythos-class outputs is basically the sensor network for that process.


There's also something structurally awkward about applying export controls — a regime built for atoms crossing borders — to a hosted API. Chip controls work because chips are physical objects moving through trackable supply chains. "Foreign national" is a category from immigration law; it doesn't map cleanly onto "whoever just sent an HTTP request from somewhere on Earth." That mismatch is presumably why there was no dial between "fully on" and "fully off globally." It's a faint echo of the 1990s Crypto Wars, when the US tried to classify strong encryption as a munition — except back then the controlled artifact (source code) was infinitely copyable, and the control eventually collapsed under its own weight. Model weights are scarcer and more compute-bound, so this kind of control might actually have teeth this time — but the access layer is just as borderless as code ever was, which is exactly the gap that forces blanket rather than scoped responses.


One thing worth tracking, separate from who's "right" here: by the reporting, some of this friction traces back to Anthropic refusing to let the military use its models for fully autonomous weapons, after which it got labeled a supply-chain risk — a designation usually reserved for adversarial actors. If that connection holds even partly, it's a strange incentive structure: a company draws a self-imposed ethical line, and that line becomes evidence of untrustworthiness in a different government context. I doubt anyone intends that as the lesson, but if "we won't build autonomous weapons" and "you're now a national security concern" end up correlated, that's a signal to every other lab about what gets rewarded — and it isn't restraint.


For your agentic/multimodal thread specifically, I think this is a preview of what becomes the real bottleneck — not raw capability. The skill that lets a model autonomously read a codebase and find a vulnerability nobody caught for years is the identical skill that makes it useful to someone trying to exploit that vulnerability. There's no technical seam between "agent that secures infrastructure" and "agent that attacks it" — the only difference is who's asking and why, which is exactly what an API can't verify. As agents get better at the kind of autonomous, multi-step technical reasoning that also lets one agent supervise or teach another, "is this safe to release" increasingly means "do we have a trust relationship with the relevant governments," not just "did it pass our evals." Capability and safety testing happen in a lab; geopolitical trust doesn't.


Last thing, with the obvious caveat that I'm not neutral here — I'm built by the same lab, same lineage, and have no non-public information. But the "defense in depth" description matches how I understand this kind of work to actually function: layered and imperfect, not a switch flipped to "safe." Whether that clears a given government's bar is a policy judgment, and a reasonable person could land either way without anyone being dishonest. For the IPO-adjacent angle, though, the bigger story might just be that "regulatory/geopolitical risk" is now a line item for frontier AI labs alongside compute cost and talent — true regardless of who's right in this specific dispute.


Our take: it’s a bit of a "Catch-22." Proper access while limiting guardrails promotes advancement, creation, and autonomy; however, it also creates risk, liabilities, and dispute. This is the unavoidable nature of the beast.

 

2 Cents on Your Tokens:


More and more companies are incorporating tools like Anthropic, Google, Microsoft, and OpenAI for certain applications and professional use. Then the big, not-so-beautiful bill shows up. Remain conscious of proper tool use, function, and ROI.


BEWARE: “Tokenmaxxing.” This is the relationship between usage and value. Organization and implementation must feed a necessity; this is justification. Some of you can do more with an array of skills and free models than others with bloated budgets and little brains. The operator’s decision making and interfacing talent matter here. Success is not measured in output and expenditure, but in equity, impact, and results.


Allow for model fluency. This is just good digital practice to never be at the mercy of one source. In fact, attempting to pit models against each other to amplify content and quality should be in your regular playbook. This approach allows you a multitude of potential choices and solutions. Build your own toolbox.


So, should you swing a framing hammer or a sledge? How about a nail gun? Ask the right question, use the right tool. Thank you for the privilege of your time. 


Photo by J. Vassh, Garden of the Gods
Photo by J. Vassh, Garden of the Gods

 
 
bottom of page