Episode 17: Measuring AI productivity, and will AI have Ads?

Published: Sunday, Jan 11, 2026 • Duration: 53 minutes • Season 1

Measuring AI productivity, and will AI have Ads?

Download MP3 | Watch on YouTube

https://x.com/kaihendry/status/2009678360046420110

https://docs.astral.sh/uv/llms.txt

Watch on YouTube

summarize "https://youtu.be/tKad1dhw5Qs" --timestamps --slides

This episode is a wide-ranging conversation that starts with a storm anecdote and quickly centers on practical questions about measuring AI-driven productivity, the human friction that accompanies adoption, how teams build trust in LLM-generated code via testing and spec-driven workflows, and the emerging business pressures (access costs, funnel collapse, and the idea of ads inside LLMs). The hosts mix concrete examples from company practice (DORA snapshots, property-based testing, runner costs) with tech-culture takes (Jevons paradox, “10x” skepticism) and tool notes (transcript indexers, conductor, spec tools).
Slide 1

Opening: storm, power and infrastructure

Conversation opens with personal storm anecdotes and a concrete resilience problem: an earlier outage left the speaker “out of power for like two 47 hours” (almost two days), exposing how quickly heating and mobile connectivity fail without electricity and prompting practical coping choices (butane cookers, spare kits) and a brief digression about unreliable mobile mast backup in rural areas. Slide 2

Execs demand numbers: “How much more productive?”

A middle manager describes being asked to quantify AI productivity gains (“how much more productive 100 100% 200%”), while customers sometimes block the best models; the practical answer discussed is to use DORA-style engineering metrics (four defined metrics from the Accelerate research), take early snapshots per team (data, cloud, app, iPad teams) and compare pre/post AI adoption using instrumented sources (GitHub, Jira, PR/release timestamps) to show measurable change. Slide 3

Metrics culture, fudging, and adoption anxiety

Speakers warn about presentation culture and gaming numbers (example: tagging coverage reported jumping from 20% to 60% after manual/retroactive edits), and about comparing adopters to non-adopters (OpenAI/other reports claim big gains but lack transparent methodology); they push back on hype labels like “10x engineer” and cite suspicious self-reports (a tool claiming it “finished it in 10% of the time”) as reasons to insist on verifiable, repeatable measurement and private challenge of fudged results. Slide 4

Spec-driven dev and trustworthy tests for LLM output

The hosts highlight a spec-driven IDE approach (Kirro) that mixes deterministic logic and formatted user stories to generate structural tests, and they recommend property-based testing (fast-check in Node, Hypothesis in Python) to validate behavior rather than surface syntax. Practical notes: porting CDK to CDKTF revealed test harness value, but also costly CI requirements (examples noted: TypeScript jobs needing ~8 CPUs and 32 GB RAM), so runner costs and caching (turbo cache, monorepo task runners) matter for real-world adoption. Slide 5

Jevons paradox, access inequality, and session tooling

They invoke Jevons paradox — “if things are getting easier to do, we’re going to do more of it, not less of it.” — to argue easier AI will expand activity, not shrink it, while flagging access stratification: a handful of heavy users spend hundreds of dollars monthly (example figure cited: people spending ~$600/month) and capture outsized output. To counter siloed learning, they demo transcript/session tooling (a Rust-based indexer and sharing tool) and argue teams need transcript rating, shared prompts, and workflows to upskill people and capture institutional knowledge. Slide 6

Workflows, conductor, and the revenue threat to docs

They discuss limits of structured LLM workflows (difficulty getting templates to do meta tasks like “library of prompts”), mixed experiences with conductor (sometimes keeps an LLM loop going), and real business risk: a docs-driven project reported an ~80% drop in funnel traffic/revenue after LLMs started generating code for users, triggering major staff reductions in one example; as a commercial response the hosts debate attribution and monetization inside LLM flows and offer a blunt hot take — ads or paid surfaced attributions inside model outputs are likely to appear — and note a viewer poll result that around 65% thought an LLM already outcodes them, summarized as “Claude is already better than them at coding”.

Model: openai/gpt-5-mini

Transcript (auto-generated from YouTube captions)
Yo, how was your trip yesterday?
>> It was pretty epic. There was a huge
storm in the UK actually and I'm and I
pretty much drove well, not through it,
but like I stopped off at my sisters and
then I and then when it passed, I drove
after it. So, I got to see the beginning
of it and the and the after of it. So,
it was a pretty epic storm.
>> Have you seen that kite surfing video?
[clears throat] There was like
>> in Israel
>> during a storm
>> there was like a thermal vortex like a
little cyclone and some kite surfers got
caught into it and one guy he went like
super high but he was able to eject when
he was lowered to the ground. The other
guy he got dragged onto rocks and died.
So it was um it's been posted like four
or five times on Reddit. Maybe because I
like to kite surfer so I see it.
>> It's pretty bad. But yeah, you're
talking about the storm reminded me.
Sorry you were saying after the storm.
Uh but I mean to be honest we were not
impacted. But one thing that makes me
nervous is that um I think last year was
it the year before we had a storm durag
and we were out of power for like two 47
hours. Is that two days almost?
And it was pretty terrible because
without power we cannot actually power
our oil
or gas heating systems.
>> Basically [snorts]
It really goes to hell here very
quickly. And uh and funny and the other
thing that really irritates me is that
like as a South African if I'm actually
used to the power to go to go out but
the infrastructure the mobile
infrastructure it always has like a huge
backup. Your audio just switched
>> from what to what?
>> I don't know the audio. Maybe when you
when you moved closer. Maybe it's it's
>> Oh, you know what? I think I I knocked
out the bloody
uh RX or something.
>> Yeah. Test.
>> No, when you said that about um you
couldn't turn on your electric like
cooking or or anything. When I wanted I
came to Vietnam, I I really hated those
gas cookers.
>> Oh, yeah. But now you love them.
>> And and the fire. I don't like it
because I see videos of these gas
canisters exploding and things like
that. And I was begging my um my wife to
get like um an electric cooker and she's
like, "No way." You know, there's going
to be a power outage and we can't cook.
>> Yeah.
>> Exactly. Exactly. Yeah. [snorts] So we
have we have like a little Korean set
and those little butane canisters for
for for these occasions.
>> But yeah, going back to the the
infrastructure cuz this is the AI
infrastructure podcast.
>> Yeah.
>> In South Africa, the infrastructure is
very robust.
But when the power goes out here in uh
you know 5 minutes later my 4G signal
disappears and then
>> really and yeah like none of the none of
the masts in England have backup power.
Maybe they do it out in like London, but
around here in the southwest, none have
backup power. [snorts]
>> You weren't afraid it was going to be
like like a Venezuela incident. They
take out the the power in the city and
next thing you know, your prime minister
is on a plane back to the US.
>> Well, yeah. Well, the resiliency factor
in in UK and I do think it's related to
the well to the mobile infrastructure is
is very poor actually. So if the if any
Russians are listening to this
>> Yeah.
>> now you know
>> but I know they take London in in u
what's this movie with um Edge with Tom
Cruz
uh where he brings tomorrow.
>> Yeah, tomorrow's edge. [gasps]
>> They take London very quickly as soon as
the the beach uh battle was lost.
[laughter]
>> Tomorrow's I watched the movie like four
or five times. I know it's maybe because
you like and you like games. Anyway,
>> the Anyway,
>> one of the things I wanted to that was
just banging around my head after being
at on on site
um the last couple of days was that
on on one on one side I have you know
since I'm a middle manager I don't know
how you feel you are in in the industry
but I mean I do feel like I'm a very
hands-on middle manager but nonetheless
I'm a middle manager since I'm my middle
manager. I have I have the the executive
sort of asking me, "Hey, Kai, um how's
this AI thing going? Um it makes people
more productive, right? It it's how much
more productive 100 100% 200% how can
you just tell can you build a case study
for me to show how how more productive
you are with AI?" And this is in that
said and the AI roller in this
particular customer is incredibly
frustrating and bad because they disable
uh all the good models. So that I got on
one hand I've got the these sort of
questions and then and then for the sort
of the colleagues sort of um that I work
with
many of them are actually um quite
nervous about the future and um okay I'm
exaggerating here when they when they
when I think in my in the in the notes I
call them doomers but like a lot of
people say like are basically like uh
saying to me that like oh um AI isn't
going to make me um you know twice as
effective. And it's absolute [ __ ]
that management are are coming up with
these things as an excuse to uh reorg
and lay off people. But that's the way
they feel and and and and
I'm in the middle and I don't know and
my feeling is is is just trying to be
well I am a positive I like to think I'm
a positive person. Some some people
might disagree with with with that
sentiment, but like I'm very excited
about the future. I am so excited, but
at the same time, like I can't really
like my my fun factor that we have in
our little podcast chats, I does that
translate to a 200% productivity? I I I
I am reluctant to say that. and and at
the same time I'm like how do I make how
do I make people more positive who are
around me or or you know
>> report to me you know what I mean
>> okay let's because you touched on like
five different points I think the first
one you mentioned is how do you me
measure productivity and that's all
about the devops accelerate report right
dora metrics how many times do you
deploy how long product feature to go to
production
>> let's be real a lot of these dora
metrics aren't aren't extremely well
recorded and and things like that and
they're very hard to
>> No, no, there are so we implemented the
Dora metrics well me uh a a colleague of
mine implemented the Dora metrics as a
previous company and you you have four
clear metrics that are well defined of
what you need to measure and they are
the key I can argue with you but carry
on
>> and like meantime between failure is is
often something that's quite contentious
I
It's you're the matter is you need to
measure where you are and then measure
the improvements, right? So it's about
taking snapshots. It's about agreeing.
Yeah, sure.
They identified these are the metrics
that correlate to productive um
companies or companies that are
successful according to the research
that they've been doing sponsored by
Google over several years looking at
many different companies and across
different industries and different
styles of companies maturity and all and
they said look those four metrics start
measuring them. So then this the fact is
source of truth. Where do I get this
information? How exactly do I measure
that for this particular team for this
particular product, right? Um where does
the the you know how long does it take
for a feature to be from specified down
to production and and how do I measure
that you know do I go to GitHub? Is
there like a certain Jira ticket
attached to the pull request and can I
see when the when the Jira is you know
progressing along the different stages
and when the is the mer is a branch
merged and the release tech and so on
right so that was the exercise really at
my previous company that um a part of
the team did and they started having
snapshots for every team the data team
the cloud team the app team you know the
iPad app team
>> snapshots by team yeah
>> yeah and then they start showing okay
this team's quite well on the way on the
on their on their journey uh after one
month two months uh improvements and so
on. So then I think if you started doing
those snapshots early before the AI
adoption you can definitely you know
compare this team has heavily
>> I mean I wish I wish it was that easy
but maybe I'm just being making coming
up with excuses but like
>> yeah it really comes down to just doing
like breaking it down in small steps and
doing it. I hate I have some colleagues
that are constantly going about what if
yes we have to understand the problem so
we can well define the solution but at
the same time if we spend days arguing
back and forth about what is that of how
it really works until like put it then
you know spin up I'm very quick to spin
up a docker compos cluster with the
different uh you know like the engineext
and whatever like this was particularly
about rate limiting or or about shared
memory right using a swap
>> rate limiting I I I worked on an enginex
rate limiting uh uh module back a couple
of gigs ago.
>> Yeah,
>> we should.
>> So, the thing is I hate I hate these
very theoretical like I want a a very
early MVP definition. What do we need to
do to to prove some of this? And that's
what the guy in in in in that comp in
the company I worked for did. He went
with like a pre-prepared
uh solution like pop flash or pop dog
something. It's not pop dog because
that's the train in in in in
Counterstrike. [laughter]
uh with it's like pop up. So something
like that. Anyway,
>> but like okay I
>> was a very clear path very didn't take
long took quick. Okay, you're talking
about making some stone cold
measurements and you're right. You're
absolutely right.
>> But I had a second point and and I think
that the objective, but like how do you
deal with people's, you know, vibes, you
know, that's
>> Yeah. So people second thing, right?
First thing is the metrics because this
is the same problem that we always face
as a platform team where we're not
directly um valued against the business
and we need to our metrics really come
down to the Dora metrics, right? we need
to optimize for those and we need to
prove that our efforts and our projects
optimize those metrics. So that's how
you can attach value to a platform team
aside from constantly asking them to to
>> So have you honestly had success? So
this is how you've defended your DevOps
uh work in the past by saying like uh
you know hey we are delivering uh faster
thanks to my work. I mean or you or you
just playing this by the book?
>> No. Um
>> do you have a real case study?
>> Yeah, but I found so the team that I was
in was doing this. I think I stayed for
6 months and then I left. [laughter] Um
because I wanted I thought there was
fundamentally something wrong with the
way that the platform was was being
developed and I wanted instead of
constantly being forced to go and you
know do little bandates on the different
teams infrastructure I wanted to have a
proper design but I've also learned that
that is not really something um but
anyway so I've seen they because there
was a very strong culture of sharing
results in that company so there was
always this big all all hands-on
meetings where every team presented the
results. So, the team presented those
dashboards, but I did think a lot of
those numbers were being fudged for
example.
>> Oh, yeah, of course. Oh my god. I was
That's what I was going to say. I mean,
if if you have a culture of doing line
graphs, it's going to be so bad.
>> Yeah. No, but also like there was this
initiative to to increase the AWS
resource tagging and one approach was to
add tags and tax policies around the
infrastructure as code so that they are
being measured and pull requests are
being uh blocked until tag tag coverage
improves and so on until beatings will
continue until the memorial improves.
[laughter] Anyway, um but in those cases
>> will continue. I've never heard that one
but it's quite hilarious. It's I think
it's from Montipython not uh you know um
yeah
>> will continue until the moral improves.
Yeah. So [laughter] but in this case
they actually went through manually and
added tags all over the place and it was
justified because some of the resources
were not under IC. But to be honest I
think it was just a a hack because they
were under IC and you just had to go put
it just to present the report like yeah
we improved tech coverage from 20% to
60%. like, "Okay, fuck."
>> I hate sitting in these meetings because
I'm just like thinking, "Oh my god, this
is bullshit." But if I say something,
then everyone's going to think I'm the
biggest [ __ ] in the room.
>> Yeah, I do say some things. But I've
learned to say those things not in
public channels anymore. [laughter]
>> But I go [snorts] to people, I say like
that's not right. You just added the
tags manually.
>> Private meetings will resume.
[laughter] So, but then the second part
you touched on is about the human aspect
like because yes sure you can measure a
team per team and you can measure those
metrics but this is something that Jean
Kim and Steve Jagger touched on in their
talk right when they said like they had
a problem in OpenAI um in where how can
you compare somebody that has fully
adopted LLM versus someone who hasn't
because if they looked at at the
traditional metrics someone who has
adopted AI is doing a lot better and and
so
>> well they say that and they don't really
tell you
>> because they sell it also, right?
>> Yeah. And this is the this is the
frustrating thing is like, oh yeah,
they're much better. They're 10 times 10
times engineer, but like uh okay, how do
you
>> that's just the 10x engineer is like a
is is a meme, isn't it? It's just means
that that that they're
>> a red flag. Once I was interview I was
interviewed for a company and I I I
passed and they offered me and I I
wasn't really accepting the offer and
then they called me in and and then the
they called in like the CTO or CEO and
he came in I heard you're like a
>> ax 10x engineer
>> or or what do you call a hero or
something like that and the moment I
hear that word and then the other
examples was like you know our
originally founders they rewrote the
whole thing in in three nights of Red
Bull Uh, and I'm like, "Yeah, no, no,
I'm not joining your company."
[laughter] If that's what you um, you
know, think is a good like work life
balance. Um, that's not the
>> living in you are living in Asia. Work
life balance.
>> Yeah. Yeah. That was
>> Here we are talking about work on on a
Saturday morning. Work life balance.
>> I'm not talking about work. I'm ranting.
[laughter]
>> This is my my my exhaust. My
>> this is my therapy. No, but like um
you're right with Jean then and um maybe
they are selling these they are selling
AI and they're selling um this. So where
are the actual measurements? I don't
know. Maybe they were referring
something that's publicly available and
I haven't read it. I just took them for
their word. So
>> Oh yeah. I mean also that that anthropic
thing I think we discussed that earlier.
I mean
>> the report the that report that was made
by a
>> Oh yeah
>> there was this entropic report and and
it was like we asked cloud to estimate
the task and then it finished it in 10%
of the time. We asked cloud to analyze
all of its work and it and it estimated
that it finished it all in 10% of the
time. It's like yeah well [laughter]
>> I guess the next thing I'm going to do
is ask ask Gemini or Claude how to u
measure the productivity gains. And of
course it does feel completely wrong to
do it that way.
>> Yeah.
>> I I mean
>> unless unless it can it can surface
hard. I'm inclined to to just pull it
back to vibes, you know, like so I was
asked to give a pres internal
presentation about the the gains of AI
and um I'm thinking of just just uh
turning turning the talk into something
like how people are happier.
uh which might not go down well and I
might get
but I think I think that's the key thing
for me. I'm actually just genuinely
happier because I can now deliver faster
and better work.
You know what a team lead of me in the
past, he used to do the standup uh you
know every weekly sync up or I think it
was weekly and he would ask the mood
like happy sad every time for every
every member to measure it.
>> Yeah. As part of the standup like you
know what did you you but it wasn't
standup maybe it was like the weekly uh
check-in or something like that. But but
team white everyone
>> Yeah. Yeah. But that's that's a typical
thing you do in a retro for example, you
know. Yeah, you have safety check.
>> If you have [clears throat] if you have
>> if you have recorded it then you can say
well since we rolled out Claude Code
plans to everyone happiness has gone up.
>> Yeah but
>> anxiousness has gone up too. [laughter]
Well, it's a difficult it's difficult to
measure that too because ultimately
>> I mean like personally
I I I can be maybe it's the my Germanic
blood but like I wake up in the morning
and I'm not freaking happy dude
>> and then people hey happy new year how's
it going like like please
>> like yeah sure
>> cloud deployed all of my features that I
I estimated for two weeks of work in
just 3 days but at the same time the
toilet on the second level is broken.
It's like [laughter]
>> Yeah. Oh, actually that that reminds me
my particular client the the office has
so many like problems.
But I do think overall I think after the
Christmas period, New Year period, the
sentiment and the way people talk about
AI has significantly changed.
People are more like declaring that they
don't look at their code.
>> Yeah. more
>> that video. I mean, you linked it to me.
I can't believe he was right.
>> Yes. From
>> that Theo video. I to be honest, I find
Theo a little bit irritating to watch
for some reason. But like I did
>> I always complain, but I always watch
his videos.
>> Yeah, I did watch his video and I
thought he I thought he was absolutely
spot on. I even I even uh
>> You're good.
>> I even tweeted. Yeah. So,
I don't know. Maybe it's his hair. But
anyway, the
>> it's just a side comments of like but I
I felt maybe I got used to them or I
feel there's less. He was very much
frustrated that Antropic was not giving
him anything and he was very much saying
OpenAI is doing a great job but maybe
because it more aligns with me now.
Maybe it's my little bubble and it fits
in my bubble better. I was like, "Oh,
he's he's hailing on Tropic now. I like
to watch his videos." [laughter] But but
I still feel like
um perhaps I need to rewatch the video,
but like um
yeah, there was that that that
[clears throat] whole spiel about not
not looking at the code anymore. But I
still feel like people
I felt like he said it well maybe I'm
just daydreaming at the moment but like
I still there's still a lot of people
who somehow throw in a weird argument
that like I'm not looking at my code but
at the same time I still I do look at I
do look at it just to make sure it's
it's producing good code like so what is
it? Are you looking or you're not
looking? You know, but I mean are are
you like what do you mean by I'm not
looking? You know, this is I feel like
people are not being super honest about
it.
>> I think it's a it's um it's a it's a
slope. How do you say it's like a a
range of of looking at code? I'm looking
at code.
>> I'm not looking uh
>> it's a range. Yes. I'm looking like at
every line or I'm looking at the
functions or I'm looking at the file
structures or I'm looking at the you
know architecture high level um you know
cleanliness of of of it perhaps um but
I had something else I wanted to say and
now it's gone oh no that the other I
thought you were going with that
sentence like I'm not looking at the
code but it's and then they say it's the
most well- tested and validated code
that I have ever delivered in my life
and and a lot of people say that like
I've never spend so much like effort on
unit tests because and I think this is
true because with the AI the unit tests
are so effective now the moment the AI
doesn't like write proper code in like
creates a small bug it gets an immediate
feedback immediately says like hey that
just broke it's like oh I forgot about
that and it fixes it and I think that's
true I think it and we and we discussed
about this for the last three podcasts
and I hope I've convinced you a little
bit but um but these
>> I'm still a little bit dubious about
even than that because in the last video
if you remember we we we observed that
the beads uh
>> no but failing we don't talk about beats
>> we don't we don't talk about beats here
I mean as as a we we talk about it as a
memory system not as a well tested
architecture because I told you in the
very early on I am very scared of
updating beats because I was just
looking at the change logs on the
releases and it was like fixing
merge conflict that corrupted all of
your issue task list I was like all the
I don't want to upgrade to [laughter]
[gasps] you know
>> I think I think DAX and open code were
were extoring their tests but I haven't
looked at open code GitHub
okay so I think so just let me talk
about
>> just to summarize our last chat
[snorts]
>> if you want to talk about the test part
because you just sent me the video of um
specdriven development from AI engineer
that was filmed. I looked at the date,
November 22, uh, from one of the
engineers that works on Kirro, the AWS
specdriven development IDE, and it was
very interesting because he dives into
exactly where do they trust or put LLM
flows and I know they because they do
property based testing and if you use
NodeJS, you can do that with fact check
or fast check. Sorry, fast check.
>> Yeah. Or hypothesis.
>> Hypothesis in Python he mentions. Yes.
>> So, I'm actually not very familiar with
this. This is so you do you know what
are you talking about then?
>> So because I was porting AWS CDK to
CDKTF for Terraform um or CDK for
Terraform there's several
um scenarios where they use property
based testing in their testing harness
that I also ported over. I think it was
in um the way that you define schedules
um for like scheduled rules or scheduled
things like if you if you if you do
cloudatch and you want to define a
schedule and you have different
constraints they do property based
testing in that part of the code and I I
never used it before and I was like what
is all this because then I have to I
have to install a fast check and I have
to learn how that works and and usually
I spend a lot of time trying to
understand how the tool works. and how
the property based testing works. But in
this case, the moment I installed it and
I copied over the to the the tests and
then I could verify that my uh Terraform
because it wasn't really testing
Terraform or cloud formation. It was
really testing the functionality of
defining a scheduled rule. So fast check
really cool. It was really nice to see
it run and I'm very excited. Um
>> okay. And so the fast check is a snap uh
sorry
>> is a NodeJS property based testing
framework. Okay,
>> I need to Okay, I want to see these
tests run in open code and I'm just
struggling to
just find out how they release.
Okay, how like I hate it when there's no
make file.
How do you do that? So one of the things
that they highlighted uh in the in the
specdriven development setup of Kirao is
they use procedural or like algorithm um
deterministic logic um to set up the
testing and the validation. So they
force you early on to define ears
formatted tests. So uh sorry um user
requirements and user stories
>> like cucumber type stuff or what are you
talking Yeah, behavior driven kind of
like
as a user when I do this I expect this
or something like that. So that they are
very strict in the in the way that they
want want the LLM to generate these and
they do a lot of like structural tests
against those because they evaluate that
and they generate property based testing
out of those markdown. So that's super
interesting, right? Because now you're
no longer asking an LLM to transcribe
the user story into some type of
testing. you're actually fundamentally
validating
um whatever the user and the LLM has
come to after you know brainstorming and
I think LLMs are great at that um help
us you know expand our horizon and think
and um and then really structurally
validate it and generate proper tests
about it u which you can trust so that
was a very nice uh aspect on it and then
the second part that he highlighted was
um in
>> I got to see this I mean do can you show
me these
Sorry, my brain is
not really working. So, do you have
maybe you can steal the screen share and
and show me these property based tests?
Do you have anything to add?
>> I didn't I don't he didn't show them in
the video. Um I would need to install
Kira and figure out where
>> Oh my god. So, we we're talking about
it. We don't need we
>> no because we are asking the question um
how can you trust the test that the LLM
is writing and the way that K is solving
it is by you know putting a lot of
formal mechanisms around it and that's
very exciting
>> this before I've never seen this before
>> yeah yeah blacksmith runs on um
depot.deb I'm actually investigating
that because I'm forking the Terraform
CDK and the biggest concern is the cost
of the runners because some of those uh
TypeScript um programs require eight
CPUs and 32 GB of memory when they are
running.
So I need to find a good uh GitHub
hosted runner alternative.
>> Okay. Do you use bun by the way? Cuz I
noticed Bun got bored.
>> Everyone's using bun now. Yeah, Antropic
acquired it. But like I'm struggling to
run this test here.
>> This is bun turbo type check.
>> Turbo is turbo repo.
>> Type check is just a package of JSON. So
turbo is a way to run uh tasks across
your monor repo and and with a cache. So
it will look at the input sha and then
if if the input sh didn't change it will
just use the cached result so it can
speed up your task. So turbo is just
going to run type check as a package of
JSON script.
>> Okay. Type check is just going to run
TSC and usually type check is just
running the TypeScript compiler. It's
not really
>> I can't work out how to run uh open code
tests.
>> Yeah.
>> Okay. Okay. If I can't work out maybe
let's have a look at the the GitHub at
least. Um Oh, sorry. I'm not sharing the
screen here.
I'm just I'm I'm
just
>> There are a lot of good things from the
video you shared me.
>> Oh, cool cool. But I need I'm a I'm a
visual learner. I need to see the bloody
test. I need I need to see the test. I
need to see it running.
I need to see the test failing. I need
to build
>> trust with the the test has value. You
know what I mean? It's like [snorts]
>> this is a journey. I maybe it's not
worth talking about now. I I kind of
want to turn the conversation back to um
Oh yeah. Like one thing that was very
good that
>> Yeah. from
>> from Theo's video is the Jven's paradox
which was actually mentioned by Simon
Willis
>> and I'm also a huge fan of Simon Willis.
>> Uh if you don't subscribe to him then
you're missing out. I really love his uh
his tool to um share uh so so this post
I wrote here is like I I use his UVX his
claude code transcripts into
>> there's another tool I wanted to show
you as well but
>> to pick out the session and then we sort
of share it as a team. So I think this
is like
>> this is like level one.
>> Yeah, this is very very level one. But
but but that's where we are at my
workplace. It's like it's like you know
the uh the different stakeholders and
the people are are basically in a
browser copy and pasting between that
and VS Code because the VS Code
organization is so locked down. Uh
anyway, so yeah paradox. Uh yeah, I'm
trying to get people I mean this is a
new argument for me and I I I'm I'm I'm
I believe it. I'm I'm behind it. It's
like if things are getting easier to do,
we're going to do more of it, not less
of it. I mean, I'm I'm I mean, I think
it's all public stuff. I have a I have a
guy that I I think I used to work with
him you um he was sort of posting some
like doomer stuff like it's like when
AGA comes along then prime real estate
is going to collapse and then I then I
asked him like why do you think that and
he says like with with heavy uh AGI job
displacement that uh people will become
on lower incomes and then people just
won't be able to to meet rent I think is
what he was trying to say and then
there'll be like a a spiral down where
everything sort of collapses
and then I do mention Jven's paradox in
that in that Twitter thread
and then he comes back by saying that's
reason by analogy. Analogy will break at
the limit.
It will hold for some time.
um as long as we're not bottlenecked by
energy supply. Uh you know, oh god. Um
but but with the cost of intelligence
labor approaching zero,
>> there'd be very few demands that supply
cannot meet. Lower incomes coupled with
lower prices would would would produce a
deflationary
cycle like unlike any we've ever seen.
Um and then then he quotes
uh this some I don't know where he gets
it exactly. We're going to hell before
everyone goes to heaven. I mean I get
that but at the same time I don't want
to get behind it. I I don't want to
think that's going to happen. I
I mean his def his his com his
combativeness
of like that Jven's paradox is going to
somehow break at the limit. I
h I don't think I I don't think so.
I mean I I don't want to think so
but um yeah regarding the the the the
fact that originally I always feel that
we are heavily subsidized they are
spending so much money and we are spend
we are actually using a lot more quote
than our quotas unless you're using API
keys which is dumb. Yeah, I mean I I
think things could get a bit more
restrictive in future. Like we're
definitely since since we're the
astronauts of of the AI, we we're
getting a free ride in some ways.
>> Yeah. Which to me feels like it's going
to become this elitists. Um only few
people can afford I think it's already a
little bit the case, right? There's only
a few people that are spending $600 a
month on three cloud max
plans and those guys are like, you know,
outputting. so much um that it's is if
like
>> yeah like Steve Jger is like elevated.
>> Yeah, he was he was he was very elevated
before with this stuff.
>> Yeah, I wanted to share something that I
said I would I would share a couple of
tools that are really cool. Um I want to
share my screen. Yeah.
>> Well, actually I'm not too sure about
this.
>> Can you see my screen?
>> Yep.
All right. So, you see here at the top I
have this little little bar that shows
like my codecs. I've not used much right
now. And I'm
>> Is that Is that a Is that an official
Codeex thing or is it some
>> No, it has Cloud here. Oh, so stupid. 5
seconds ago, Cloud was right there. But
I guess it I ran it on on uh another
one.
>> I'm not I'm not a huge fan of menu bars.
I always feel like they're just sucking
power. But [snorts] you find this one
useful.
>> Well, what did you
>> for me from what I do right now? Because
you kept asking how do you keep track of
the context and how do you keep track of
the usage? But I I don't care so much
about context because you get a warning
about 30% in. So I'm like, okay, that's
when I kill the session.
>> I probably have to cut this out. You I
think you you accidentally shared
someone's name there.
>> Where?
>> On the bottom left is a this
>> Oh, this one.
>> Yeah. No, he it's a book for about um AI
changes tools, but you still owe the
craft. I like that that uh page.
>> Okay. I'm just a bit nervous about
having to edit this video.
>> It's not a CV or anything. It's a book.
>> Oh, okay.
>> Um so I think you'll be happy book.
>> Oh, that's the author, I guess. Okay.
>> Yeah, he's the author and he works at
shopback in Singapore and he's kind of
as an evangelist talking a lot about the
tools because he's introducing them. I
think he's bit of an engineer.
>> So, is he is he bumping up AI? Ah, well,
this is the kind of
>> I don't know how how how we should
reason about it. Yeah. But the fact is
get back to
>> maybe share that with me if it's if it's
good cuz I need the reason it's open on
my on my laptop is I still haven't read
it.
>> All right. Right. But but just just just
to summarize where I am, Vincent, I need
to
>> to basically um become an AI mascot in
my company without
>> Yeah, these are definitely the recovery
codes you should cut out.
>> [laughter]
>> Thanks.
Uh, okay. Hold it. Do I I need to really
cut that out. Okay, let me just make a
so I can find it. Cut.
>> Your name escort. You can start from
there.
>> Okay. What were you showing me? What
were you showing me? You showing me
>> Hold on.
>> Oh, you were showing me this like
codeex. uh you're showing me this menu
bar thing
you're you're flexing to the the podcast
listeners or pe the viewers like yeah
I've got I'm burnt I'm burnt through a
billion tokens in the last 24 hours
stuff that you wouldn't couldn't afford
or you couldn't get through your company
no that's where we were we were talking
about how it's like an elitist right but
that made me realize I have this little
bar here at the top that shows you how
much codeex how cloud and then how much
Gemini I've used so far.
>> How much of the Earth you've destroyed?
>> Yeah, that's So people are doing these
Instagram
>> $500.
>> No,
>> that's API calls for sure because I only
pay $20 a month. So So just to show that
if you pay API calls, it's a bit silly
because they cost a lot of money. But if
you use something like um a $20
subscription, you get that in mind.
But still actually these these
>> but this is this is this is wrong
because this is probably because I
haven't spent extra usage 500
>> as a consultant
people people it's it's sometimes eye
watering but it's not unusual to see
people being build you know thousand
$1,000 US a day right for a consultant
and and then and then I at the same time
I get a bit you know I I shout out when
I see someone that has spent $600 on
Claude in a month. Of course,
in real terms, AI is incredibly more
efficient than a than a consultant.
[snorts] Okay, maybe I have to cut that
out.
Yeah, I'm trying to find So, this is um
cloud cuz you were showing a way to
share cloud sessions. You can take a
session and you can create a nice web
page from it. Yeah. Using Simon Willis's
tool.
>> Yeah, I can link to the my blog.
>> And then there's this thing called the
the cloud um I don't know what it stands
for like assistant session reviewer and
it's Cass
and it you run an indexer across your
cloud session directory and it will it's
using oh it's using Rust so it must be
fast, right? Okay.
>> Oh, come on. This is trivial stuff, but
you could do this in shell.
>> No, but the fact is that that like
>> Okay. So, so you're basically going
across all your conversations.
>> Yeah. Across everything and you can type
like a small little like a single word
and it will find all of them. I think
it's doing a fuzzy search. I'm not sure.
But my indexer has an error apparent on
its another problem is
>> so so when it comes to these transcripts
I feel like um there needs to be a tool
to say like
um to to rate to rate the transcripts
right like like this was a really good
session this this session sucked what
went wrong here what went right here you
know these these sort of like metadata
so we have a way of sharing transcripts
but we don't have a great way
>> oh look at even works across open code
and codex in Gemini.
>> Yeah. Nice. So we don't have we don't
have a nice way of of really
learning as a human how to interact with
AI. I feel because if we if we feed
these transcripts back into AI and
asking how to improve I think that I'm
not too sure it might work but
>> yeah but that's not
>> but I'm I'm talking about upskilling
teams getting getting people more to
build trust with the and building shared
prompts and things like that.
>> Yeah. Yeah. Yeah. This is what I'm this
is what I'm trying to do at work.
>> So, this is one thing that they
highlight in um somebody asked a
question in the KO talk saying, "Can I
customize the templates and the prompts
and the workflow and they're like they
are over here?" I don't remember the
answer, but it it seemed like given the
way that they are so structured, I think
it's hard for you to change those
things. And um I definitely feel that
those templates, those workflows, like I
try to use speckit to set up some type
of uh ADR. So it's funny because speckit
is ADR in a way kind not really exactly
the same, but it's in a way you're
capturing like a decision
record, right? Um, so I was trying to
ask specit to create a decision record
and it couldn't really get it because
like, oh, you're building a program. So,
and I was like, yeah, no, but we're
using LLMs. I was like, oh, so you're
building a program that um is going to,
you know, you know, you're going to call
out to the open API and and I was like,
no, no, no. I'm trying to do something.
[laughter] So, it was really uh really
difficult to get it to do kind of like
meta meta spec type of things like I
want you to create a library of prompts
to use for my community planning
sessions and things like that. It was
like huh I only know how to create a
program just like so you want me to
create a program to run your community
meetings? No, no, that's not what I
want. So, it was um so yes, you need to
be able to modify these workflows. Uh I
tried conductor first by the way to have
a feel of it.
>> Oh, yeah. Oh,
>> did you have a good experience like I
did?
>> No, actually both conductor even though
in the in the product design and
definition I said very clearly like what
we're building here is a way to like a a
type of a process for the repository. I
want to build a pro. So it's a repo for
keeping track of decisions and proposals
and spec and so on and and um no didn't
go for that. Again, probably not the
right project for it. But yeah, other
than that, conductor was pretty cool. I
I feel it actually kept kept Gemini
going like uh I thought at some point
when I finished with all of the upfront
planning, Gemini went in a loop and for
a significant amount of time kept going.
I've never seen Gemini keep going like
that uh in a while. Have you had the
same experience with conductor like
that? It can keep keep Gemini going a
bit like
>> Well, it was a good experience. So I
guess I I don't I mean it was a week ago
if not more. So I guess I guess it was
doing its thing. Yeah, I guess it was
doing it thing. [snorts] Oh, okay. You
know what? I need the weather is
actually good right now. It's not
raining. I feel like I need to wash my
car and run around while there's some
sunlight. So I think I'm going to end
the conversation a bit early if you
don't mind. I still think I still think
what we talked about was valuable. I
don't know if if we I guess at this
point I would like to engage anyone
who's listening to please comment below
and uh chime in with your thoughts
around making AI uh I mean are you
positive about AI and things like that
or another interesting thing that Theo
did he had like a little poll on his
video to say that do you think Opus or
do you think Claude is outputting better
code than you are at this point and I
think it was like 65%
agree agreed that Claude is already
better than them at coding, which I
thought was interesting. And to be
honest, I agree. Like I don't know how
you feel. Me too.
>> Yeah, me too. I think it's better
because it's better than me.
>> Uh it's more like it has a broader
knowledge base like it can produce code
for a broader set of things that I don't
know.
>> Yeah. But even if I ask it to produce Go
code or or code or languages I know
quite well, I I don't I think it's
pretty okay.
>> No, but I liked Steve's take saying that
it's way better at Go than at
Typescript. I always felt TypeScript is
perfect because it gives you the
structure and the schemas, but Steve
says types is TypeScript is way too you
can do way too fancy things with types
and TypeScript and Go is just really
down to the you know straightforward. So
so AI is better with Go. There there's
some arguments made I think oh I can't I
can't remember his name now uh the
educator guy he was who makes some
typescript courses he was arguing that
types and AI go well together because
you want those that that that test
>> I always feel like that but Steve Jger
was saying that that he he he rewrote
gas
>> typescript was was sucks in that regard
>> he says it's it's a token waste it's
it's generating a ton of tokens that
don't really do much and uh and with
>> code is so much cleaner
and direct although that I like what I
like with go is this very um you know
write everything twice mentality of go
which is like don't over generalize no
generics for ages don't like just
rewrite the code rewrite the functions
>> I still I still don't even quite
understand generics I mean I don't use
them so maybe I'm missing out
>> I I did originally when I was writing
code golang but when I asked the llam to
write I I go like yeah don't like I
don't mind it being duplicated in
several places cuz I'm like
>> Yeah, exactly. It's cheap. Yeah. To do
all the different uh
>> you need to maintain it. It's going to
figure out
>> switches and all that sort of stuff.
Yeah. Yeah. I get it. I mean, go it. I
mean, I'm feeling pretty pretty uh #
uh what do you call it? Inspired #
grateful. because I know Go better than
most people and uh and I know AI better
than most people in some ways and
everything everything's coming together
really nicely for me. I feel and I'm
touching wood here. is like everything's
going well except maybe my pay packet.
But uh
you also did one final uh because the
whole tailwind laying off 75% of staff
and then people saying well well if your
business model was like what was their
business model right how that they got
disrupted. I don't know if you know
about this whole Tailwind scenario.
>> Oh yeah I didn't I did notice that was
that was interesting. Uh I think I
didn't read the whole post and he also
made a podcast himself, right? Where
he's walking his
>> I read the post and I listened to the
whole podcast.
>> Oh, really? So you probably have more
details, but but my my uh what do you
call it? Headline reading interpretation
was that was that um yeah the LLM was
was was taking their business away
because uh you didn't need to use
Tailwind and its abstractions I suppose
to do to work with CSS. You can just get
Claude to do it.
>> Yeah.
>> And then it boiled down also to the
website. So I think there's this llm.ext
and they removed it because they don't
want to
>> the thing that triggered it is somebody
went I'm going to here's a pull request
to add the llm.txt to your docs.
>> And he was I'm sorry but I have to
decline this. I just had to lay out 75%
because the docs are our only funnel and
nobody's reading them because the LLMs
are just generating code. Nobody's
reading our docs anymore. And it's how
we we funnel our business. We've got 80%
down uh in revenue. Um for you know was
it 80% down in traffic through through
uh those funnels. So he's measuring
those different you know ways that they
are driving business right and and a lot
of discussions
>> stack overflow problem in a way but
whatever it's different. Yeah. Uh so a
lot of people were saying like from
Tailwind they realized as a as a an
expert um I have a network and people
come to me to and and that's how I get
jobs but that all falls away with AI. Uh
and and then some some of the uh um he
says I've been sponsoring Tailwind for a
long time. Love what they do. I learned
a lot from them. So in his video was
more like highlighting all the good work
that they've done and he shared like
there's this even if you are using an
LLM reading their design um you know
guides like they have they written a
book on how to do the UX which is not
just about the CSS but really about like
how to build uh UX. He says that's a
really good book then you should really
read it and he's also encouraging people
to sponsor them.
>> Yeah.
>> Yeah. But that I mean I'm sponsoring is
always not a solution right? You need a
viable business business model. If their
business model is upselling through
their docs, yeah, they're in problem. I
mean, and interestingly, Charlie from
the Astro Project mentioned that that
they have LLM's text, but they also have
a similar they're they're in the similar
boat
>> because they need to upsell their other
um
>> they're hosting, I guess. Oh, they have
some um
what is there? I can't even remember
what they're business. [snorts]
>> There's a there's a there's a product
that I think they were working on. Oh,
this Pyx thing.
>> I think that's
>> going to be uh a paid for thing.
>> That's like Dino then. Dino and Bun like
reinventing the packaging. Well, not
Bun, but Dino reinventing the packaging.
So, so you get the point that everything
could just Yeah, they're in the similar
boat. Okay. Anyway, I'm I'm interrupting
I'm interrupting you.
>> No, the other thing that Tio said is as
a as a business, if you let LLM generate
your code, go check the dependencies and
see if you can go sponsor them like if
but but to me it seems like the real
solution is is ads, right? This is how
Warp they say like we're giving you Opus
4 access for a fraction of the cost
because they put
>> no one's going to read the docs ads. How
you
>> Yeah. Because ultimately inside the docs
are ads. Like if you like to if you like
this here's our our offering, right? So
if you translate it, nobody's going to
the dogs. It means that those ads no
longer live with the dogs, but they must
be surfaced by the LM somehow, right?
And Warp's uh operating model is like we
give cheap AI access in our terminal by
serving you ads and that's like
and and you know if OpenAI would
introduce that you know the alterate is
already created but it's the only way
like if OpenAI can semantically
>> any business that requires ads now I
guess with with LLMs in the mix LLMs are
filtering out the ads
>> so any business model that relies on ads
is basically in trouble here and it
probably includes includes Google.
>> Google of course. Yeah, Google. Google
had a big re wakeup call. They had all
hands on deck
>> uh early in the in the Chad GPT saga.
>> So what is what is Google's response
there
>> to become a leader because they have put
so much money. They have their own TPUs.
They build their own Gemini CLI. They
they become an LLM platform.
>> But like okay, go going back to
Tailwind. I mean don't say sponsorship
is the answer because it's not. What is
Vincent? How do you think they're gonna
come out of this or how how is Astral
going to come out of this?
>> I think if the LLMs I guess it's it's
the whole problem with LLM's not giving
attribution, right? If there's a way for
LLMs to give better attribution and it's
like I'm you know imagine that while
it's generating your website it goes and
by the way there's a sale on Tailwind
right now. [laughter]
I'm I'm using Tailwind to build your
website. No. Um,
>> that sounds like another weird form of
advertising which sounds insane.
>> It would be funny if if if you're try if
you're using the MCP or you're
troubleshooting the issue and you're
going like, you know, it's still not
looking right. There's still an
overlaying bug. It just stops. I just
said, go to the Tailwind team, dude.
Like, [laughter] go ask them to fix to
help you fix it.
>> Pay them. Um, cuz I don't have the
answer here. [laughter]
>> Yeah.
>> Oh god. I think if a company is smart
and and onropic is smart because they
did like the double limits usage limits
for for the whole Christmas holiday that
was a gigantic success. Look at how many
people tried it out. They were like I
got so much done and I'm 100% in on this
right now. I actually upgraded to $100.
>> They just they're like drug dealers,
aren't they?
>> Yeah. So if there's if if Antropic is
listening, Boris, [laughter]
>> if they're smart, what they're going to
do is they're going to make
>> more correct.
>> Yeah. like loop detected. Um, please
like based on the context it looks like
like Clippy pops up. Looks like you're
trying to fix a wind a tailwind issue.
[laughter]
Here's the tailwind guys website.
>> No, I think your proposal I mean it's
frankly
>> ads are coming. Ads are coming. I think
this is
>> there's got to be a better Okay,
>> I'm going to do a hot take. Hot take
2026 hot take from Vincent. Ads in your
LLMs this year. Guaranteed warp's
already there. So I already what? Not
except for warp. Let's talk.
>> Arguably there already ads like when you
run claude it gives you like a tip. Try
this mobile experience. Try this.
>> But it's always there. It's always like
use our our designer plugin for better
web design. What about it goes like use
the Tailwind MCP subscription only $5 a
month. [laughter]
Okay. Bye. Go wash your car.
Let's see how many people hate comments.
[snorts]
>> Okay, please please please rate the
podcast. Please comment below. Please
tell me how we tell us how we can
improve. Actually, I did get feedback
from from from Swix, but the the
Riverside thing looks a bit too
expensive for me at least. It's like £24
a month
>> to to improve the quality. I was like, I
don't want to spend money.
>> Everything's improving except my pay
packet.
Uh, I did earn I Oh, yeah. I should I
should note maybe I can give you access
to my YouTube studio, but I think I I've
earned two pound
thanks to this whole podcast series. So,
thanks Vincent for coming on the show.
>> That's two pound off for a family of
four.
>> Two pound. I I'll I think I'll can buy
you a coffee when I see you in Vietnam
or something.
>> Yeah. Come to Vietnam. You can feast
>> for exactly one meal.
>> Wow. [laughter]
Yeah, I that's probably the real issue
that we should have talked about is how
people's pay because Madu Shan who I was
chatting to on X, he was right to point
out that that people's income is not
improving.
>> Okay.
>> Yeah. [clears throat]
>> Okay. See you guys.
>> See you.
>> See you guys. Oh my god. I say you guys.
Sorry. So, see you Vincent.
>> Comment below. Bye.
>> You guys and girls. Come on.
>> Oh yeah. Sorry.