@tokentrove avatar
@tokentrove

Philip Terry

Ai, Blockchain, Marketing, Music, ActorUnited States of America

6Following 14Followers
Posts
Pages

Recent posts

For the last few years, the AI conversation has mostly been about models. Which one wins the latest benchmark. Using AI looks like asking a question & getting an answer, then starting over. Working with AI is different. It means the system remembers your life, work, & increasingly decides which version of itself is right for the job. That shift is already showing up in 3 places at once. 1st Autonomy. AI companies are training to actually do a job: use software, make decisions, fail, recover, try again. Cursor introduced a router that reads a request and decides which model should handle it. xAI launched Grok Bot, that picks its own model for a task. The model isn't the product anymore. The decision about which model to use is. 2nd Memory. Google's Gemini can now pull from your all your info to answer based on your actual life, not a generic one. For years, every AI conversation started from 0. That's no longer true, & it's easy to underestimate how big a change that is. The 3rd piece is where all of this runs. Meta just released Glimmer, its Muse model small enough to run on a normal laptop, no cloud required. That sounds like a footnote. Once a system knows this much about you & can act without asking, where it runs becomes a question of trust. Some of this work, people want happening on their own machine. Put those 3 together. AI is becoming more autonomous, more personal, & more local, all at once. Knowing how to write a clear prompt is useful. It's no longer the hard part. The harder skill is knowing what work should be handed to AI, what context it needs, when it should act on its own versus when a person stays involved, & what you're comfortable letting it see. That's a different kind of literacy. Not a world where people simply know how to use AI, but one where people understand how AI works, understand how they work, & build something in the space between the two. The next few years are about AI & people together.
Two years ago the entire AI conversation was about Token Maxxing, chasing the newest model and pushing as many tokens through the system as possible. I have written about this before. What's happening right now is the next chapter, and it looks nothing like the last one. The question is no longer which model is biggest. It's which model earns its token count. Kimi K3 landed this month and it is worth watching. It runs $3 input and $15 output per million tokens, roughly 40 percent cheaper than Claude Opus 4.8. It is not a clean sweep, Claude still leads on the harder production coding work, but K3 wins outright on frontend generation and several agentic tests. Cost and capability used to move in the same direction. They are starting to split. OpenAI is telling the same story from a different angle. GPT-5.6 now ships in three tiers instead of one: Sol for frontier reasoning, Terra as the balanced daily workhorse, Luna for fast, high-volume tasks. Spinning up a flagship model to reformat a script is like flying a commercial jet to the grocery store. The tier system exists because not every prompt deserves the same horsepower. Here is where it gets genuinely interesting. Cursor just launched something called Cursor Router, live for a few days now. It is not a new model. It is a decision layer that reads every request and sends it to whichever model actually fits the job, claiming frontier quality at 60 percent lower cost. Auto mode inside Cursor runs on it by default now. That is the industry admitting, that intelligent routing is infrastructure now. It is also worth asking who owns that decision layer. Cursor is mid acquisition by SpaceX in a deal built to strengthen Grok. Whoever controls the router controls where the token spend actually flows, and that is a bigger strategic asset than the market seems to be pricing in right now. There is a deeper flywheel story underneath this, Tesla's data, Starlink, SpaceX's valuation. That's a post on its own. #AI #LLMs #TokenEfficiency #Cursor #OpenAI #KimiK3
Last week I was cutting the grass in a heat wave and it hit me: the thing most people are missing about AI's next decade was sitting right there. I'd been listening to an interview with Emad Mostaque. His argument: economics has always been about managing scarcity, and intelligence has always been the scarcest resource. As intelligence gets closer to costing nothing, that foundation cracks. He calls it the Intelligence Inversion, a shift from scarcity to abundance. I realized abundance isn't free. It has a physical floor. Right now, that floor is heat. Every chip running these models generates heat. There's a hard ceiling on how much you can pull out of a rack of silicon before the system throttles or fails. That ceiling has become the defining constraint in AI infrastructure. Not the algorithms. The thermodynamics. For years, the AI story was about training, the massive compute to build a model. But training happens once or twice. Inference, running the model every time someone sends a prompt or an agent takes an action, happens millions of times a day. Inference is quietly becoming the bigger cost. That shift is why a company called Etched is worth watching. They are betting that a chip built for 1 job, inference, could beat Nvidia's general-purpose GPUs. Their chip, Sohu, runs at a fraction of the power of a typical AI chip. Whether the bet pays off is unclear. A specialized chip is faster at its one job, but riskier if the model architecture shifts. OpenAI joined AMD, Broadcom, Meta, Microsoft, and Nvidia this year in a consortium built to standardize the shift from copper to light based connections in AI data centers. As AI agents take on longer, more autonomous work instead of single prompts, inference demand compounds. Every interaction needs power and generates heat. The abundant intelligence future, depends on solving that constraint at scale, with light: fiber optics. We flesh suits aren't running out of ideas about what abundant intelligence could do. We might run out of ways to keep it cool.
The day SpaceX went public, my phone lit up. Texts from family. Calls from friends. "Phil, should I buy SpaceX?" The excitement was real. The urgency was real. What most of them didn't realize is that anyone with a retirement plan likely already has a small SpaceX allocation sitting in their portfolio right now, thanks to the accelerated market entry the stock received. That dynamic is exactly the lens I want to bring to this week's news. SpaceX went public and reportedly acquired Cursor, the AI coding platform, for $60 billion. The commentary has been loud and mostly focused on valuations and personalities. But if you filter the noise and look at what's actually being assembled underneath, something structural and quiet is happening. This isn't a rocket company buying a software tool. SpaceX lowers the cost of putting computing hardware into orbit. Starlink provides borderless internet with low latency to every corner of the planet. xAI builds frontier models. And now Cursor embeds those models directly where software gets written. Not a chatbot window. The development environment itself. So what piece is getting almost no attention: AI agents don't have bank accounts. When an autonomous coding agent needs to rent GPU space to compute, pull an external dataset, or commission a specialized agent to audit its work, it cannot wait three business days for an ACH transfer to clear. It needs programmable, instant value exchange. Tokenized payment rails. That's why the stablecoin regulatory framework moving through Washington right now isn't just a crypto story. It's the financial layer of this same stack being laid in real time. The workforce of the next decade won't be measured in headcount. It will be measured in how well humans architect, govern, and add judgment to the agents doing the work. I'm not here to tell you what to buy or who to root for. What I am saying is that the plumbing of the global economy is being replaced while most people are still asking whether they should buy the pipes.
Earlier this week I talked about the "weird middle" of AI adoption and why most people are still using the word "agent" to describe something much closer to a really smart template. Here's where it gets more interesting. True AI agents, the kind that actually observe environments, select tools, and take autonomous action toward a goal, won't just process text and summarize documents. Eventually they'll negotiate. They'll transact. They'll book your travel, renegotiate your subscriptions, and manage purchases on your behalf in real time. And when they do, they'll need to move money. Not in three business days. Not through a legacy bank account waiting on an ACH transfer. Instantly. Programmatically. Without a human approving each transaction in the middle of the chain. That's exactly why what happened in Washington a couple weeks ago matters more to the AI conversation than almost anyone is saying right now. The CLARITY Act cleared a major milestone in the Senate, advancing a bipartisan framework to regulate payment stablecoins in the United States. A stablecoin is essentially a digital dollar backed 1:1 by liquid reserves, a dollar that can move at the speed of code. Most of the coverage has framed this as a crypto story. It isn't. It's an AI infrastructure story. The financial rails that true agents will need to operate autonomously are being laid right now, while most people are still debating whether Siri's new voice sounds more human. The plumbing is going in before most people know a house is being built. Reality Check: AI and digital value exchange are not parallel conversations. They are converging. The companies and people who see that now will read the next wave clearly. Everyone else will wake up surprised. We are not just building smarter software. We are building a new operating system for how value moves in the world. How many people in your network are connecting these two conversations?
"Oh! Phil I made an agent at work." My family said that at dinner last week, genuinely excited. I smiled and nodded because I love that she's engaging with this stuff. But in my head I was doing the translation work I always do now. She hadn't built an agent. She'd built a really smart template. And that distinction is quietly becoming one of the most important gaps in how people understand where this technology is actually going. That same week, six different professionals told me they use Microsoft Copilot every single day at work. When I asked if they knew it could connect directly to ChatGPT under the hood, they all looked at me like I'd spoken another language. To them it's just a helpful corporate utility. Something the company installed. They're using the technology. They're just not in a relationship with it yet. This is the "weird middle" of AI adoption. And Apple's WWDC announcements this week made it impossible to ignore. The next generation of Siri and Apple Intelligence aren't just feature updates. They're the moment AI gets woven directly into the fabric of daily life, your photos, your emails, your personal context. The moment it stops being a tab you open and starts being infrastructure you breathe. Reality check: we've started using the same words to describe completely different things. A workflow follows steps you define. A GPT follows instructions you provide. A true AI agent observes its environment, makes independent decisions, selects its own tools, and takes action toward a goal with minimal human oversight. It doesn't pause to ask permission at every step. It determines its own next move. That's a different species entirely. The "weird middle" is where mainstream is just arriving at conversations builders were having two years ago. That gap is frustrating if you're waiting for the world to catch up. It's a gift if you're paying attention. What tech term do you hear getting misused most right now?
Yesterday I called out Token Maxing: the 2026 version of busy work where you run circles with prompts and feel productive while your actual outcomes stay flat. So how do you actually Output Max? It's not a magic prompt. It's a system. Here's the framework I use. Start with the friction, not the tool. Most people open Gemini and ask "what can this do?" Wrong first question. The right question is: where is my workflow leaking energy? Is it the two hours you spend synthesizing research? Is it manually moving notes from a meeting into your project board? Output Maxing starts with a real problem, not a feature list. Build the conveyor belt. Once you have the problem, don't just throw a chatbot at it. Map the flow first (Mermaid.ai). Chain the tasks (Google Workflows). Connect the pieces (Google Stitch). Instead of running 10 prompts manually, you build something where data flows in, gets processed, and lands on your desk ready for a final judgment call. That's a system. That's Output Maxing. The Taste Test This is where the real work lives. Once the output arrives, you don't just hit send. You ask: does this sound like a human or a machine trying to pass as one? Does this solve the client's actual problem, or just the one they asked for? Is it grounded? AI provides the scale. You provide the soul. Here's the reality: Output Maxing isn't about working more with AI. It's about working less on "stuff" so you can spend your 3 billion seconds of experience on the things that actually matter. Which step in your workflow is leaking the most energy right now? #AI #SystemsThinking #FutureOfWork #OutputMaxing #HumanFirst P.S. The framework only works if you're honest about where the friction actually lives. Most people already know the answer. They just haven't stopped moving long enough to look at it.
I will be making this next post a two parter for tomorrow! Token Maxing is the 2026 version of looking busy. I'm watching teams run into the same wall right now. Chasing the newest model. Testing the latest prompt. The screen is moving, but business is still stuck. There is a difference between AI activity and AI output. Token Maxing is tool-focused. It's running 50 prompts to get one mediocre blog post. It's opening 12 tabs just to summarize a meeting. It feels like progress because your screen is full. Output Maxing is outcome-focused. It's the person who uses Claude, Gemini, Grok ect. to eliminate three steps of a process entirely so they can go get a coffee and actually think. AI Isn’t Just a “Tool”. It’s a System. Most people treat AI like a hammer. If you have a hammer, you look for nails. But AI is closer to the plumbing system for your business. You don't "use" plumbing. You design your house around it so water flows where it needs to without you thinking about it. If you focus on what AI can do, you'll keep adding tools until you're overwhelmed. If you focus on what needs to get done, you'll start building systems that compound your effort. Those are two completely different games. The "Human" Reality Check I'm not a cheerleader for the software. I'm a cheerleader for the humans trying to navigate this without losing their minds. The winners of this decade won't be the ones with the most tokens used. They'll be the ones who kept their agency, their judgment, and their strategic taste while letting the agents handle the heavy lifting. Stop collecting tools. Start refining outcomes. Tomorrow I'll share the actual framework I use to move from Token Maxing to Output Maxing. How many AI tabs do you have open right now? #AI #FutureOfWork #SystemsThinking #TechLeadership #OutputMaxing

Connect with Philip Terry

Join INSPIRED