· Dan Shipper

Dan Shipper: The AI Paradox — More Automation, More Humans, More Work

As AI automates more, humans do more work, not less — every agent needs a human, so the winning stance is to be simultaneously AI-pilled and bullish on humans, and to ride each new model into whatever you do.

aifuture-of-workagentssaasautomationproduct-managementdesignhiring0% confidence

Why this is in the corpus

Dan Shipper (CEO of Every) runs perhaps the most AI-forward company in operation and has an unusual track record of calling paradigm shifts early (Claude Code / Cowork a year ahead). This episode is a dense, forward-looking set of contrarian predictions about how work, SaaS, and roles change over the coming year — signal-rich for operators deciding where to place bets.

Summary for skimmers

Shipper argues the AI job apocalypse is overstated: models make yesterday's competence cheap and commoditized, pushing humans to a new frontier. Work bifurcates into a company super-agent (in Slack) and a Codex/Cowork work surface where humans and agents collaborate via an in-app browser. He is bullish on SaaS (agents increase usage; users bring their own tokens), on PMs and full-stack designers, and on the forward-deployed engineer role. The through-line: every agent needs a human, so automation creates work rather than removing it.

Briefing

What survives the editorial filter

This page should feel like a smart colleague already listened for you and left only the operating logic worth keeping. Not everything said in the episode makes it through.

Trust signal

Direct episode extraction

Best used for

Decision-grade retrieval metadata not yet added for this episode.

Hold lightly

No explicit downgrade reason stored yet for this episode.

Principles

Durable claims that survive beyond the speaker's biography — each with explicit limits, transferability judgment, and evidence.

Principle

Models make yesterday''s human competence cheap and commoditized

Each model release commoditizes yesterday''s human competence, so durable value moves to the frontier where humans turn cheap capability into something new.

Because everyone uses the same models, default output all looks the same — slop tweets, identical landing pages. The financial incentive to keep models compliant and aligned means they structurally trail the humans pushing past them.

Do not compete on what the model already does — compete on what you build on top of it.

what models do in general is they make yesterday''s human competence cheap and so it becomes commoditized. It''s not valuable anymore. What humans do is we go in there and we''re like, yeah, we, we have all this frozen human competence from yesterday. How do I use this? Like make something new and interesting.Dan Shipper

Principle

Automation is a lie — every agent needs a human on top of it

Every automation you deploy requires a human on top of it to keep it working, so automation creates supervisory work rather than eliminating labor.

Shipper frames automation as management: managers are not on the beach, they check in constantly. The same holds for model management — it takes real time and attention. This is why an AI-forward company still hires humans.

Budget human attention for every agent you deploy — the labor moves, it does not disappear.

Automation is a lie in the sense that every time you automate something in order to make sure the automation is working well, you need a human on top of it.Dan Shipper

Principle

Benchmarks rise only on problems we can frame and score

Benchmark saturation measures only framed, scorable work, so it never equals full human replacement — the framing itself stays human.

Shipper notes he could trivially rewrite his own benchmark to zero out the newest model, because there is always a higher frame to move to. Benchmark progress is real but partial.

Distrust the leap from "benchmark saturated" to "job automated" — the unframed work is where humans stay.

benchmarks rise on problems that we''ve framed that we can articulate, that we can score. And there''s a lot of work that''s human work that it, it can''t be scored until you write it down, but the act of thinking to prompt it or write it down is, is something that you can''t measureDan Shipper

Principle

Ride the models — using each new release is the durable job-security move

To stay employed through AI progress, ride the models: apply each new release to whatever you do.

Riding the models is not one fixed action because they keep changing — it is staying curious and playful, applying every new model to your job or life and re-testing what it can now do.

Make trying each new model against your real work a standing habit.

The only thing you need to do is ride the models and that means use them for whatever it is that you do.Dan Shipper

Principle

Predict the future by living in it, not by prognosticating

Foresight comes from building a pocket of the future you live in and then noticing, not from abstract prediction.

Every staffs entirely early adopters and reviews models, giving them alpha/beta access and a lived vantage point. Writing about what they notice both crystallizes it and makes it real for others.

Stop forecasting; build a team that lives in the future and report what you see.

what you don''t wanna do is prognosticate what do you, what you wanna do instead is, is just live in it together.Dan Shipper

Principle

The edge of AI is wherever it meets a real human doing something

The frontier of useful AI is at the point where a model meets a real human task, not where models are built.

Shipper argues Every in Brooklyn is "quite far ahead of people in San Francisco" on usage, because whoever applies a new model first discovers what it is good for — a genuine discovery available to anyone.

Being first to apply a new model to a real task is a form of discovery open to everyone.

I think the edge of AI is wherever AI meets like a real human doing something because the people in San Francisco, they''re making it, but they don''t actually know a lot about How to use it.Dan Shipper

Principle

A company only goes as far in AI as its CEO does — it is not delegable

AI fluency at the top is a hard ceiling on the whole company — the CEO cannot delegate their way to intuition.

It currently looks like a CEO can get away with an unchanged day, but Shipper predicts that reverses rapidly into "I''m way behind." Senior leadership hands-on use is the differentiator.

If you lead, put your own hands in the tools — you cannot delegate the intuition.

your company''s only gonna go as far as your CEO goes in AI and it''s not something you can delegate. You have to have your hands in it ''cause you don''t, otherwise you don''t have an intuition for it.Dan Shipper

Principle

AI writing is good if you stand behind every line; slop is what you don''t

Judge AI-generated work not by its origin but by whether the author stands behind every line — that is the slop line.

Shipper welcomes AI-generated documents but bans sending one you cannot discuss; a well-directed GPT-5.5 strategy doc beats most hand-written ones because the bar for human strategy writing is low.

Send AI-assisted work only if you can defend every line of it.

there''s a difference between an AI generated document that''s slop and not, and the slap one is it took them less time to make it than it takes me to read it. And they don''t stand behind every line.Dan Shipper

Principle

Generalists can go much further now, especially at small companies

The AI era rewards generalists, who can now execute across functions that used to need specialists — a boon for small companies.

Every deliberately hires generalists who love touching many areas. Shipper expects roles to re-settle over time (marketing people still do marketing), but generalist reach is a durable new advantage.

Hire and cultivate generalists — the tools now let one person cover ground that took several.

I also think that you can get a lot further being a generalist now, and that''s like really cool, especially for, for smaller companies.Dan Shipper

Principle

The human and the agent are now on the same piece of work together

Work is shifting from delegating to agents to collaborating with them live on the same surface, with mutual visibility.

Shipper''s Codex-plus-Proof setup embodies this: Codex watches him write, he watches Codex, both act in one place. Software built only for human use or only for agent use misses this middle.

Build tools where human and agent see each other and act together in real time.

we''re moving into this new paradigm I think where the human and the agent are on the same piece of work together and they''re both doing things and you need to have, I need to have visibility into what the agent is doing. The agent has to have visibility into what I''m doing.Dan Shipper

Principle

Be simultaneously AI-pilled and bullish on humans

The winning stance is to be maximally AI-adopting and maximally human-bullish at the same time.

This resolves the paradox of an AI-forward company doubling headcount: heavy automation increases, not decreases, the need for human judgment and supervision.

Reject the framing that AI enthusiasm and hiring humans are in tension.

I''m simultaneously extremely AI pilled extremely and very bullish on humans and the role of humans in making sure that AI is working well.Dan Shipper

Principle

Agents increase the number of SaaS users, they do not replace SaaS

Agents multiply SaaS demand instead of destroying it, because agents become high-volume users of the same tools.

Shipper reports Every''s own SaaS spend is up year over year despite heavy internal agent use, and would "buy SaaS stocks right now." The SaaS apocalypse is "dumb."

Do not short SaaS on AI fears — agents are new customers.

I think that what agents do is increase the number of users of SaaS, not get rid of it. And so I think SaaS companies are going to see like an insane spike in the amount of demandDan Shipper

Frameworks

Reusable systems and operating models — including when they help and when they break.

Framework

The Allocation Economy — working with AI is a manager''s job

Working with AI is best modeled as management, and the diagnostic is that managing takes real, continuous time — not beach time.

Diagnostic test: ask whether you picture the AI-augmented human on the beach or checking in constantly. The benchmark discourse makes AI look more autonomous than it is; the allocation-economy lens corrects for that by importing what managers actually do all day.

Plan for managing agents to consume time and attention, the way managing people does.

I wrote this piece a couple years ago called the allocation about the al allocation economy. Like the idea that the, the way that humans are gonna work with AI is gonna, is gonna be like, like being a manager. And the thing that you have to remember about managers is like managers actually spend a lot of time working. Most managers are not like on the beachDan Shipper

Framework

Build your own private benchmark to measure real model capability

Construct a private benchmark from a real artifact and human-expert rewrites, then design prompts that reveal capability without leaking the answer.

Diagnostic: does your prompt let the model demonstrate agency, or does it hand it the plan? Shipper''s naive "here are the issues, go fix them" prompt made every model paper over the codebase; only a first-principles rewrite prompt revealed that GPT-5.5 would actually rip out and rewrite bad code (scoring 62 vs ~30 for prior models, vs high-80s/low-90s for human seniors).

Build a private eval on your real artifacts; iterate the prompt until it exposes true capability.

So I made this senior, it''s called a, the senior engineer benchmark. And it''s like, how good is AI versus a human engineer? And the way that I built it is again, have this app proof, I just vibe coded it on the sideDan Shipper

Framework

Two-mode work bifurcation — a company super-agent plus a Codex/Cowork work surface

Work with agents bifurcates into an async company super-agent you delegate to and a synchronous Codex/Cowork surface where your daily work happens.

Diagnostic axes: is the task delegated-and-await (super-agent, Slack) or co-present (work surface with in-app browser)? Shipper expects super-agents to start at the top of the company and trickle down into specialized team/personal agents as models get less fiddly.

Sort your workflows into "delegate to the company agent" vs "do it with me on my work surface."

it''s going to bifurcate in this in two main ways how you, how you use agents. One is you''re going to be doing... everyone''s gonna have at least in their company, at least one agent that they talk to that can do work that they can offload work to... second is that most of the work that you do is actually going to happen on your computer in an environment like Codex or Claude CoworkDan Shipper

Signals

What appears to be shifting, for whom it matters, and what happens if you ignore it.

Signal

Within a year, work bifurcates into a company super-agent plus a Codex/Cowork OS

Shipper predicts that within roughly a year, most work will run through a company super-agent plus a Codex/Cowork work surface — and asks to be scored on it in May 2027.

He hedges on exact timing but commits that it should be "not obviously wrong" within a year and moving in that direction. This is the episode''s spine prediction, explicitly set up for a one-year scorecard.

Watch for company super-agents and agent-native work surfaces to go mainstream within a year.

most of the stuff that I''m ta I''m gonna talk about will be pretty apparent within a year, but it, it probably, it may, it may take longer than that... it objective, it should within at least a year be like not obviously wrong.Dan Shipper

Signal

Pull-request volume skyrockets as non-technical people ship code

A higher share of the company (ops, consulting, editors) now ships code, so PR volume skyrockets and the bottleneck moves to review and coherence.

Shipper cites Open Claw''s Pete getting thousands of PRs a day, spinning up 50,000 Codex instances to sort them and merging ~1,000. The new scarce question is not "can we build it" but "does it fit the coherent whole — and what do we delete."

Staff for the review-and-coherence bottleneck, not the build bottleneck.

the number of pull requests that you get is like skyrockets. You know, we have people, you know, in consulting or in ops roles or whatever who are or, or editors just like making pull requestsDan Shipper

Signal

CLIs are over — GUIs return as the main surface for agent work

Shipper predicts the CLI era for agent work is ending and GUIs return, with most technical staff at Every already off the CLI as their main surface.

He clarifies CLIs will not vanish (they persist for decades) but the "CLI is the reason it works" belief was wrong — GUIs deliver the same benefits, and programmers now mostly flip into Codex/Claude Code/Cursor GUIs.

Do not over-invest in CLI-only workflows — the GUI wave is coming back.

CLIs are over, we we speed ran the CLI era. It was nice while it lasted but I think it''s pretty, it''s pretty clear it''s not that cli... I would estimate that definitely the majority of the technical people inside of every are not using CLI anymore as their main work surface.Dan Shipper

Signal

Buy SaaS stocks — the SaaS apocalypse is dumb, demand will spike

Shipper''s contrarian market call: buy SaaS stocks now, expecting them up majorly over the next couple years as the apocalypse thesis fails.

Grounded in two mechanisms from the conversation: agents multiply SaaS users, and users bringing their own AI tokens restores SaaS margins. He caveats "not investment advice."

Treat the "SaaS is dead" consensus as a contrarian long, per Shipper.

I would buy SaaS stocks right now. I would, I think the SaaS apocalypse is dumb and SAS stocks will be up majorly in the next couple years. Not not investment advice, but you know, I would buy SAS stocks.Dan Shipper

Signal

GPT-5.5 jumped the senior-engineer benchmark to 62; senior-level in a year or less

On Shipper''s private benchmark, GPT-5.5 leapt to 62 (prior models ~30, humans high-80s/low-90s), implying senior-engineer-level coding within a year or less.

What set GPT-5.5 apart was agency: only it had the "sense of agency and confidence" to rip out old code and rewrite from first principles (using an Opus 4.7 plan), rather than papering over the edges.

Expect senior-engineer-grade coding models within a year — plan roles around framing, not typing.

all the models until GPT 5.5 got like a 30 out of a hundred and senior, like a human senior engineer gets like high eighties, low nineties out of a hundred. So there''s a lot to go. And then I tried GPT 5.5 and it got like a 62... It''s like very, it''s very clear that in a year or less it''s gonna be senior engineer level, right?Dan Shipper

Opportunities

Only included where there is a buyer, a real wedge, and a plausible revenue path — not vague idea theater.

Opportunity

Forward-deployed engineer as a service

The forward-deployed engineer is a real, growing role, and lending that capability out as consulting is a market others want.

Every runs an internal agent (Claudie) for its consulting practice and rents FDE capability to clients. The TAM logic: as agents proliferate, so does demand for the humans who manage them — automation "created one or many" jobs, not fewer.

Consider productizing agent-operations/FDE expertise as a service; demand grows with agent adoption.

the whole four deployed engineer concept I think is for real. And it comes out of, every agent needs a human... We also do consulting. So we, we, we lend that out to people and, and I think that''s a big, that''s a big thing that, that people want.Dan Shipper

Opportunity

Bring-your-own-tokens SaaS restores software margins

Building SaaS that humans and agents use together, where users bring their own AI tokens, restores margins that AI-inside products erode.

Instead of building AI into the product as the primary surface (which costs tokens), make a piece of software humans and AI want to collaborate on — harder to build but cheaper to run and, once built, "a good business."

Bet on BYO-token SaaS over vendor-funded inference as the durable model.

with proof, for example, anyone who uses it, I don''t pay for tokens because they''re just bringing, they bring their AI to to proof. And so it changes what you build as a SaaS company and you build it now for both humans and agents to use at the same time and it changes your margins back to, well I don''t really have to pay for tokens anymore ''cause the user''s gonna bring ai.Dan Shipper

Opportunity

Software built for human-and-agent collaboration (approval inbox, logs, rollback)

There is an open product category for software designed for concurrent human-plus-agent work: approval inboxes, change summaries, logs, and instant rollback at agent scale.

Shipper notes agent-native products can be simpler (the agent handles formatting) yet need new affordances and infrastructure — GitHub is "having problems right now" precisely because agent usage is skyrocketing. New tools can start leaner and faster than legacy incumbents.

There is greenfield in agent-native collaboration tooling — approvals, logs, rollback, agent-scale infra.

how you display that to the user is gonna be very different... you need like approval, you need a sort of inbox that sort of summarizes here''s all the stuff that''s going to happen or has happened. You need, you need logs and the ability to roll it back real quick.Dan Shipper

Lessons still worth keeping

Useful takeaways that did not fully clear the bar for durable principle status.

Lesson

A one-line hunch typed into Codex sourced the perfect L&D hire

Shipper typed a vague candidate hunch into Codex, walked away, and it returned the ideal L&D hire — whom he then DMed and met for dinner.

This is his favorite example of async agent research: work that "would''ve taken so long before." He generalizes it to sales sourcing and recruiting as high-value use cases.

Hand fuzzy sourcing/research hunches to an async agent and let it run while you work.

I feel like someone who is into, who, who who worked at General Assembly and is now into AI would be really good. And I just like literally typed it into Codex and then like went off and was doing something else and I came back and it found like this the perfect guy. It was like worked at General Assembly, was an instructor like is super AI pilled and follows me on Twitter.Dan Shipper

Lesson

Vibe-coded Proof collapsed on launch; two independent senior rewrites became the benchmark

Shipper''s vibe-coded Proof kept crashing post-launch and Codex could not fix it; he brought in two senior engineers whose independent rewrites became his senior-engineer benchmark.

The incident cost him publicly ("egg on my face"), sleep, and even bursitis from over-coding. The resolution — two human rewrites — gave him the human baseline against which he now scores every new model.

Do not ship vibe-coded systems no human on the team can debug.

the day after launch, it was like just every like 10 minutes the servers would go down and people were looking at me and I''d be like, I don''t know what''s going on. Like Codex fix it. And Codex was like, I don''t know what''s going on. Or really Codex was like, I do know what''s, what''s going on. I fixed it and then it, it would cause four other errorsDan Shipper

Lesson

Codex sent an investor email unreviewed — and it was exactly what he''d have written

Codex accidentally sent an unreviewed investor email that turned out to be exactly what Shipper would have written, showing how much routine writing is safely delegable.

Shipper now has most of his email written by GPT-5.5 and Codex and would prefer to label it as such. He still wants to decide the substance; the sentences "don''t matter that much" for rote mail.

Delegate the drafting of prosaic messages; reserve your judgment for what they should say.

I had to send an email to, to one of our investors and I asked Codex like, go do it. And us like Codex knows to ask me and it usually does, but this time it didn''t and it just sent the email and I didn''t look at it at all and I was like, fuck. And so I went to my sent and looked at it and I was like, oh, this is exactly what I would''ve sent.Dan Shipper

The Plays

Try these this week

Verb-first executable actions — each one tied to a stated outcome in the episode.

Build a private benchmark from your own artifact and human-expert rewrites

Outcome: Turn a real failure into a private benchmark: capture human-expert rewrites as the baseline, then score each new model with a capability-revealing prompt.

Context: Shipper''s naive prompt ("here are the issues, go fix them") made every model paper over the codebase; the prompt that worked ("this is vibe-coded slop, rewrite from first principles") exposed GPT-5.5''s willingness to actually rewrite. He notes he could re-baseline to zero out any new model, keeping the bar honest.

I got a, I got actually two different senior engineers to fix it independently. So I have two different rewrites of the code base that tells me how they did it, right. And so what I get to do is when we get new models, I just give the new model a prompt.
Dan Shipper
Set up once; re-score per model drop per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

  7. 7

Before you start

  • · A real artifact you own
  • · Access to human experts for the baseline
  • · Access to each new model to re-score
engineering

Run quarterly planning through Notion agents that interview every employee

Outcome: Run quarterly planning by having one strategy-seeded agent interview each employee and generate each team''s plan, then review for dependencies and quality.

Context: Shipper fed the agent the top-level company strategy; it asked each person what happened last year, their goals, metrics, and how it related to the company idea, then produced "incredibly good" per-team quarterly plans. He then identified which teams needed to talk to each other and which plans were high vs low quality.

when we did our, our quarterly planning for every, at the end of 2025, we did it all with notion agents and we just had a bunch of notion agents and or we had really one notion agent and then we had a top level company strategy and then we had everybody in the company just talked to an agent and it asked them about what happened last year, how did it go, what were your goals?
Dan Shipper
One planning cycle (quarterly) per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

  7. 7

Before you start

  • · A documented company strategy
  • · An agent platform (e.g. Notion agents)
  • · Employee participation
operations

Only serve users who arrive via an agent — skip onboarding entirely

Outcome: Draw a hard line that you only serve users who arrive via Codex/Cowork, so their agent supplies context and setup instead of an onboarding flow.

Context: Learned from Every''s hosted Open Claw product (whose waitlist they had to pause because the harness was too hard to maintain): instead of building a Slack workflow or web form asking who you are and your goals, the user pastes a prompt and their agent hands over "all the stuff I''ve been working on with Dan," yielding a custom experience — and can be told "go fix it" when something breaks.

if instead you you just, you just make a hard line of we are only going to service users who use Codex or Cowork. What happens is you just paste something into, you just paste a prompt into Codex or cowork, it goes and talks to the app and the app can be either just a regular server or, or it can be its own agent and Codex has so much information about you that it can just give it
Dan Shipper
Applies from first touch; ongoing for support per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

  7. 7

Before you start

  • · An agent-accessible interface (server or agent)
  • · Acceptance of a narrower agent-using user base
  • · Immediate reflection of agent actions in the human UI
product

Deploy one company super-agent owned by a forward-deployed engineer, then trickle down

Outcome: Start with a single company super-agent owned by a forward-deployed engineer, then let specialized team/personal agents grow downward as models mature.

Context: Shipper flipped from believing in personal-agent-per-person (too much maintenance, breaks constantly) to the one-agent-per-company model used by Shopify (one), Ramp (one), and model companies. Often the super-agent handles one job everyone needs (e.g. data requests) before specializing.

you basically set up a four deployed engineer or someone with that sort of profile who''s responsible for making sure that that agent is working for the whole company. And then maybe you have some, like some little team agents... it starts top at, at the top and then it sort of starts to trickle down where you may get more specialized agents
Dan Shipper
Top-down over months as models mature per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

  7. 7

Before you start

  • · A forward-deployed engineer or equivalent
  • · A shared surface such as Slack
  • · Willingness to maintain the agent continuously
operations

Run one Codex thread per project with an in-app browser watching you work

Outcome: Maintain one Codex thread per project and use its in-app browser so the agent watches and works alongside you on the actual document.

Context: Shipper reports 10 straight days at inbox zero using this: Codex gathers all email via Cora, renders a page, and he monologues at each email ("go research this," "collect four years of documents"). "All the stuff that I would procrastinate on I don''t really procrastinate on anymore."

I just go into one of my, one of my Codex threads, which I have one thread for every project and I just open the in-app browser, I go to the document, I usually do it in Proof... And then I just have Codex running and watching me in Proof and Codex can see what I''m doing, I can see what Codex is doing. It''s all kind of in one place.
Dan Shipper
Continuous daily use; 10 straight days at inbox zero reported per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

  7. 7

Before you start

  • · An agent harness with an in-app browser
  • · A markdown/web app the agent can see and act in
  • · Willingness to work in a GUI rather than raw CLI
productivity

Turn over the rock — re-test hard tasks with every new model

Outcome: Keep re-testing previously-impossible tasks against each new model release — the rock you turned over before may now reveal a solution.

Context: This is Shipper''s concrete method for "riding the models": be curious and playful, apply each new model to whatever you care about, and keep turning over rocks because it may not work now but probably will eventually. It generalizes beyond work.

when a new model comes out, I like always turn the rock over again to be like, can I do it now? You know, so it, you know, it could not do the senior engineer benchmark last time and I turned it over, turned the rock over again and now it''s at a 60 out of a hundred
Dan Shipper
Ongoing; per model drop per
  1. 1

  2. 2

  3. 3

  4. 4

  5. 5

  6. 6

Before you start

  • · Access to the latest models (possibly on personal time)
  • · A list of previously-impossible tasks
  • · A curious, playful mindset
productivity

Decision Moments

Actual decisions, real outcomes

Specific decisions narrated in the episode with their outcomes and transferable lessons.

Every is arguably the most AI-forward company in operation — everyone uses Codex, Cowork, and Claude Code — and the market expected such a company to need fewer humans as models improved.

Did: Doubled headcount from ~15 to ~30 people over the year while going deeper on AI, hiring engineers, designers, writers, editors, salespeople, and customer service — all early adopters.Outcome: Company grew to ~30 while remaining AI-native; Shipper frames it as proof that heavy automation increases the need for humans (every agent needs a human), resolving the automation paradox in favor of hiring.

An AI-forward company doubling headcount is not a contradiction — automation creates supervisory and framing work, so AI adoption and hiring compound.

Part of an emerging decision pattern across multiple episodes

When Open Claw launched, everyone at Every adopted it and Shipper was convinced the future was a personal agent per person — a parallel shadow org chart of little reflections of each employee.

Did: Reversed the bet: abandoned personal-agent-per-person after seeing the harnesses break constantly and demand too much maintenance, and moved to one company super-agent owned by a forward-deployed engineer (mirroring Shopify and Ramp).Outcome: Every and peer companies converged on the one-agent-per-company model; Shipper still expects personal agents to return once models are independent enough to not require fumbling with internals.

Because agents currently need a human who cares about them, centralize a maintained super-agent now and decentralize later as reliability improves.

Part of an emerging decision pattern across multiple episodes

Shipper had been the loudest public advocate for Claude Code, but OpenAI shipped GPT-5.3, a general-purpose Codex model, and the Codex desktop app that leapfrogged the Claude Code-to-Cowork paradigm.

Did: Made Codex his daily driver, spending nearly all his time in it and only occasionally flipping to Claude, on the judgment that OpenAI had "gotten back the mandate of heaven" and was getting the human-plus-agent paradigm right — while explicitly keeping the choice reversible as a horse race.Outcome: Codex became his primary work surface (email, docs, analytics via in-app browser); he remains willing to switch back if Anthropic retakes the lead, and still uses Claude a lot.

Ride the models means switching daily-driver tools on evidence, not loyalty — hold tool choice loosely in a fast horse race.

Tensions surfaced

Contradictions and trade-offs the episode raises — judgment calls a thoughtful operator has to navigate.

Tension

Everything has changed and nothing has changed

Both are true: every role has transformed while the structure of work (SaaS, jobs, Slack, email) persists — it is just another horizon.

Shipper cautions against vivid future-stories that feel real in the moment but prove too simple; the discipline is to withhold sweeping narratives until you can actually see it by living in it.

Resist both utopian and apocalyptic AI narratives; expect transformed roles inside a familiar structure.

somewhere it''s sort of a both everything''s changed and nothing has. And once you get there, I think you''re, you''re sort of starting to see like, oh yeah, this is a real thing.Dan Shipper

Tension

So much automation, yet working way more — AI-pilled vs bullish on humans

More automation produces more human work, not less, because every agent needs a human — so being AI-pilled and human-bullish is one coherent position.

Shipper lives the paradox: an AI-forward company that doubled headcount, with him working more despite heavy automation. The resolution is the allocation economy plus "automation is a lie."

Do not treat AI adoption and human hiring as a trade-off; they compound.

I''ve been feeling the, like we have so much automation, so much ai and I also work way more.Dan Shipper

Corpus connection

Where this episode fits for retrieval

What kinds of decisions this briefing is best pulled into.

Primary decisions

  • strategic-bet
  • hire
  • build-vs-buy
  • tool-adoption