Episode 17: Measuring AI productivity, and will AI have Ads?
Published: Sunday, Jan 11, 2026 • Duration: 53 minutes • Season 1
Download MP3 | Watch on YouTube
https://x.com/kaihendry/status/2009678360046420110
https://docs.astral.sh/uv/llms.txt
summarize "https://youtu.be/tKad1dhw5Qs" --timestamps --slides
This episode is a wide-ranging conversation that starts with a storm anecdote and quickly centers on practical questions about measuring AI-driven productivity, the human friction that accompanies adoption, how teams build trust in LLM-generated code via testing and spec-driven workflows, and the emerging business pressures (access costs, funnel collapse, and the idea of ads inside LLMs). The hosts mix concrete examples from company practice (DORA snapshots, property-based testing, runner costs) with tech-culture takes (Jevons paradox, “10x” skepticism) and tool notes (transcript indexers, conductor, spec tools).

Opening: storm, power and infrastructure
Conversation opens with personal storm anecdotes and a concrete resilience problem: an earlier outage left the speaker “out of power for like two 47 hours” (almost two days), exposing how quickly heating and mobile connectivity fail without electricity and prompting practical coping choices (butane cookers, spare kits) and a brief digression about unreliable mobile mast backup in rural areas.

Execs demand numbers: “How much more productive?”
A middle manager describes being asked to quantify AI productivity gains (“how much more productive 100 100% 200%”), while customers sometimes block the best models; the practical answer discussed is to use DORA-style engineering metrics (four defined metrics from the Accelerate research), take early snapshots per team (data, cloud, app, iPad teams) and compare pre/post AI adoption using instrumented sources (GitHub, Jira, PR/release timestamps) to show measurable change.

Metrics culture, fudging, and adoption anxiety
Speakers warn about presentation culture and gaming numbers (example: tagging coverage reported jumping from 20% to 60% after manual/retroactive edits), and about comparing adopters to non-adopters (OpenAI/other reports claim big gains but lack transparent methodology); they push back on hype labels like “10x engineer” and cite suspicious self-reports (a tool claiming it “finished it in 10% of the time”) as reasons to insist on verifiable, repeatable measurement and private challenge of fudged results.

Spec-driven dev and trustworthy tests for LLM output
The hosts highlight a spec-driven IDE approach (Kirro) that mixes deterministic logic and formatted user stories to generate structural tests, and they recommend property-based testing (fast-check in Node, Hypothesis in Python) to validate behavior rather than surface syntax. Practical notes: porting CDK to CDKTF revealed test harness value, but also costly CI requirements (examples noted: TypeScript jobs needing ~8 CPUs and 32 GB RAM), so runner costs and caching (turbo cache, monorepo task runners) matter for real-world adoption.

Jevons paradox, access inequality, and session tooling
They invoke Jevons paradox — “if things are getting easier to do, we’re going to do more of it, not less of it.” — to argue easier AI will expand activity, not shrink it, while flagging access stratification: a handful of heavy users spend hundreds of dollars monthly (example figure cited: people spending ~$600/month) and capture outsized output. To counter siloed learning, they demo transcript/session tooling (a Rust-based indexer and sharing tool) and argue teams need transcript rating, shared prompts, and workflows to upskill people and capture institutional knowledge.

Workflows, conductor, and the revenue threat to docs
They discuss limits of structured LLM workflows (difficulty getting templates to do meta tasks like “library of prompts”), mixed experiences with conductor (sometimes keeps an LLM loop going), and real business risk: a docs-driven project reported an ~80% drop in funnel traffic/revenue after LLMs started generating code for users, triggering major staff reductions in one example; as a commercial response the hosts debate attribution and monetization inside LLM flows and offer a blunt hot take — ads or paid surfaced attributions inside model outputs are likely to appear — and note a viewer poll result that around 65% thought an LLM already outcodes them, summarized as “Claude is already better than them at coding”.
Model: openai/gpt-5-mini
Transcript (auto-generated from YouTube captions)
Yo, how was your trip yesterday? >> It was pretty epic. There was a huge storm in the UK actually and I'm and I pretty much drove well, not through it, but like I stopped off at my sisters and then I and then when it passed, I drove after it. So, I got to see the beginning of it and the and the after of it. So, it was a pretty epic storm. >> Have you seen that kite surfing video? [clears throat] There was like >> in Israel >> during a storm >> there was like a thermal vortex like a little cyclone and some kite surfers got caught into it and one guy he went like super high but he was able to eject when he was lowered to the ground. The other guy he got dragged onto rocks and died. So it was um it's been posted like four or five times on Reddit. Maybe because I like to kite surfer so I see it. >> It's pretty bad. But yeah, you're talking about the storm reminded me. Sorry you were saying after the storm. Uh but I mean to be honest we were not impacted. But one thing that makes me nervous is that um I think last year was it the year before we had a storm durag and we were out of power for like two 47 hours. Is that two days almost? And it was pretty terrible because without power we cannot actually power our oil or gas heating systems. >> Basically [snorts] It really goes to hell here very quickly. And uh and funny and the other thing that really irritates me is that like as a South African if I'm actually used to the power to go to go out but the infrastructure the mobile infrastructure it always has like a huge backup. Your audio just switched >> from what to what? >> I don't know the audio. Maybe when you when you moved closer. Maybe it's it's >> Oh, you know what? I think I I knocked out the bloody uh RX or something. >> Yeah. Test. >> No, when you said that about um you couldn't turn on your electric like cooking or or anything. When I wanted I came to Vietnam, I I really hated those gas cookers. >> Oh, yeah. But now you love them. >> And and the fire. I don't like it because I see videos of these gas canisters exploding and things like that. And I was begging my um my wife to get like um an electric cooker and she's like, "No way." You know, there's going to be a power outage and we can't cook. >> Yeah. >> Exactly. Exactly. Yeah. [snorts] So we have we have like a little Korean set and those little butane canisters for for for these occasions. >> But yeah, going back to the the infrastructure cuz this is the AI infrastructure podcast. >> Yeah. >> In South Africa, the infrastructure is very robust. But when the power goes out here in uh you know 5 minutes later my 4G signal disappears and then >> really and yeah like none of the none of the masts in England have backup power. Maybe they do it out in like London, but around here in the southwest, none have backup power. [snorts] >> You weren't afraid it was going to be like like a Venezuela incident. They take out the the power in the city and next thing you know, your prime minister is on a plane back to the US. >> Well, yeah. Well, the resiliency factor in in UK and I do think it's related to the well to the mobile infrastructure is is very poor actually. So if the if any Russians are listening to this >> Yeah. >> now you know >> but I know they take London in in u what's this movie with um Edge with Tom Cruz uh where he brings tomorrow. >> Yeah, tomorrow's edge. [gasps] >> They take London very quickly as soon as the the beach uh battle was lost. [laughter] >> Tomorrow's I watched the movie like four or five times. I know it's maybe because you like and you like games. Anyway, >> the Anyway, >> one of the things I wanted to that was just banging around my head after being at on on site um the last couple of days was that on on one on one side I have you know since I'm a middle manager I don't know how you feel you are in in the industry but I mean I do feel like I'm a very hands-on middle manager but nonetheless I'm a middle manager since I'm my middle manager. I have I have the the executive sort of asking me, "Hey, Kai, um how's this AI thing going? Um it makes people more productive, right? It it's how much more productive 100 100% 200% how can you just tell can you build a case study for me to show how how more productive you are with AI?" And this is in that said and the AI roller in this particular customer is incredibly frustrating and bad because they disable uh all the good models. So that I got on one hand I've got the these sort of questions and then and then for the sort of the colleagues sort of um that I work with many of them are actually um quite nervous about the future and um okay I'm exaggerating here when they when they when I think in my in the in the notes I call them doomers but like a lot of people say like are basically like uh saying to me that like oh um AI isn't going to make me um you know twice as effective. And it's absolute [ __ ] that management are are coming up with these things as an excuse to uh reorg and lay off people. But that's the way they feel and and and and I'm in the middle and I don't know and my feeling is is is just trying to be well I am a positive I like to think I'm a positive person. Some some people might disagree with with with that sentiment, but like I'm very excited about the future. I am so excited, but at the same time, like I can't really like my my fun factor that we have in our little podcast chats, I does that translate to a 200% productivity? I I I I am reluctant to say that. and and at the same time I'm like how do I make how do I make people more positive who are around me or or you know >> report to me you know what I mean >> okay let's because you touched on like five different points I think the first one you mentioned is how do you me measure productivity and that's all about the devops accelerate report right dora metrics how many times do you deploy how long product feature to go to production >> let's be real a lot of these dora metrics aren't aren't extremely well recorded and and things like that and they're very hard to >> No, no, there are so we implemented the Dora metrics well me uh a a colleague of mine implemented the Dora metrics as a previous company and you you have four clear metrics that are well defined of what you need to measure and they are the key I can argue with you but carry on >> and like meantime between failure is is often something that's quite contentious I It's you're the matter is you need to measure where you are and then measure the improvements, right? So it's about taking snapshots. It's about agreeing. Yeah, sure. They identified these are the metrics that correlate to productive um companies or companies that are successful according to the research that they've been doing sponsored by Google over several years looking at many different companies and across different industries and different styles of companies maturity and all and they said look those four metrics start measuring them. So then this the fact is source of truth. Where do I get this information? How exactly do I measure that for this particular team for this particular product, right? Um where does the the you know how long does it take for a feature to be from specified down to production and and how do I measure that you know do I go to GitHub? Is there like a certain Jira ticket attached to the pull request and can I see when the when the Jira is you know progressing along the different stages and when the is the mer is a branch merged and the release tech and so on right so that was the exercise really at my previous company that um a part of the team did and they started having snapshots for every team the data team the cloud team the app team you know the iPad app team >> snapshots by team yeah >> yeah and then they start showing okay this team's quite well on the way on the on their on their journey uh after one month two months uh improvements and so on. So then I think if you started doing those snapshots early before the AI adoption you can definitely you know compare this team has heavily >> I mean I wish I wish it was that easy but maybe I'm just being making coming up with excuses but like >> yeah it really comes down to just doing like breaking it down in small steps and doing it. I hate I have some colleagues that are constantly going about what if yes we have to understand the problem so we can well define the solution but at the same time if we spend days arguing back and forth about what is that of how it really works until like put it then you know spin up I'm very quick to spin up a docker compos cluster with the different uh you know like the engineext and whatever like this was particularly about rate limiting or or about shared memory right using a swap >> rate limiting I I I worked on an enginex rate limiting uh uh module back a couple of gigs ago. >> Yeah, >> we should. >> So, the thing is I hate I hate these very theoretical like I want a a very early MVP definition. What do we need to do to to prove some of this? And that's what the guy in in in in that comp in the company I worked for did. He went with like a pre-prepared uh solution like pop flash or pop dog something. It's not pop dog because that's the train in in in in Counterstrike. [laughter] uh with it's like pop up. So something like that. Anyway, >> but like okay I >> was a very clear path very didn't take long took quick. Okay, you're talking about making some stone cold measurements and you're right. You're absolutely right. >> But I had a second point and and I think that the objective, but like how do you deal with people's, you know, vibes, you know, that's >> Yeah. So people second thing, right? First thing is the metrics because this is the same problem that we always face as a platform team where we're not directly um valued against the business and we need to our metrics really come down to the Dora metrics, right? we need to optimize for those and we need to prove that our efforts and our projects optimize those metrics. So that's how you can attach value to a platform team aside from constantly asking them to to >> So have you honestly had success? So this is how you've defended your DevOps uh work in the past by saying like uh you know hey we are delivering uh faster thanks to my work. I mean or you or you just playing this by the book? >> No. Um >> do you have a real case study? >> Yeah, but I found so the team that I was in was doing this. I think I stayed for 6 months and then I left. [laughter] Um because I wanted I thought there was fundamentally something wrong with the way that the platform was was being developed and I wanted instead of constantly being forced to go and you know do little bandates on the different teams infrastructure I wanted to have a proper design but I've also learned that that is not really something um but anyway so I've seen they because there was a very strong culture of sharing results in that company so there was always this big all all hands-on meetings where every team presented the results. So, the team presented those dashboards, but I did think a lot of those numbers were being fudged for example. >> Oh, yeah, of course. Oh my god. I was That's what I was going to say. I mean, if if you have a culture of doing line graphs, it's going to be so bad. >> Yeah. No, but also like there was this initiative to to increase the AWS resource tagging and one approach was to add tags and tax policies around the infrastructure as code so that they are being measured and pull requests are being uh blocked until tag tag coverage improves and so on until beatings will continue until the memorial improves. [laughter] Anyway, um but in those cases >> will continue. I've never heard that one but it's quite hilarious. It's I think it's from Montipython not uh you know um yeah >> will continue until the moral improves. Yeah. So [laughter] but in this case they actually went through manually and added tags all over the place and it was justified because some of the resources were not under IC. But to be honest I think it was just a a hack because they were under IC and you just had to go put it just to present the report like yeah we improved tech coverage from 20% to 60%. like, "Okay, fuck." >> I hate sitting in these meetings because I'm just like thinking, "Oh my god, this is bullshit." But if I say something, then everyone's going to think I'm the biggest [ __ ] in the room. >> Yeah, I do say some things. But I've learned to say those things not in public channels anymore. [laughter] >> But I go [snorts] to people, I say like that's not right. You just added the tags manually. >> Private meetings will resume. [laughter] So, but then the second part you touched on is about the human aspect like because yes sure you can measure a team per team and you can measure those metrics but this is something that Jean Kim and Steve Jagger touched on in their talk right when they said like they had a problem in OpenAI um in where how can you compare somebody that has fully adopted LLM versus someone who hasn't because if they looked at at the traditional metrics someone who has adopted AI is doing a lot better and and so >> well they say that and they don't really tell you >> because they sell it also, right? >> Yeah. And this is the this is the frustrating thing is like, oh yeah, they're much better. They're 10 times 10 times engineer, but like uh okay, how do you >> that's just the 10x engineer is like a is is a meme, isn't it? It's just means that that that they're >> a red flag. Once I was interview I was interviewed for a company and I I I passed and they offered me and I I wasn't really accepting the offer and then they called me in and and then the they called in like the CTO or CEO and he came in I heard you're like a >> ax 10x engineer >> or or what do you call a hero or something like that and the moment I hear that word and then the other examples was like you know our originally founders they rewrote the whole thing in in three nights of Red Bull Uh, and I'm like, "Yeah, no, no, I'm not joining your company." [laughter] If that's what you um, you know, think is a good like work life balance. Um, that's not the >> living in you are living in Asia. Work life balance. >> Yeah. Yeah. That was >> Here we are talking about work on on a Saturday morning. Work life balance. >> I'm not talking about work. I'm ranting. [laughter] >> This is my my my exhaust. My >> this is my therapy. No, but like um you're right with Jean then and um maybe they are selling these they are selling AI and they're selling um this. So where are the actual measurements? I don't know. Maybe they were referring something that's publicly available and I haven't read it. I just took them for their word. So >> Oh yeah. I mean also that that anthropic thing I think we discussed that earlier. I mean >> the report the that report that was made by a >> Oh yeah >> there was this entropic report and and it was like we asked cloud to estimate the task and then it finished it in 10% of the time. We asked cloud to analyze all of its work and it and it estimated that it finished it all in 10% of the time. It's like yeah well [laughter] >> I guess the next thing I'm going to do is ask ask Gemini or Claude how to u measure the productivity gains. And of course it does feel completely wrong to do it that way. >> Yeah. >> I I mean >> unless unless it can it can surface hard. I'm inclined to to just pull it back to vibes, you know, like so I was asked to give a pres internal presentation about the the gains of AI and um I'm thinking of just just uh turning turning the talk into something like how people are happier. uh which might not go down well and I might get but I think I think that's the key thing for me. I'm actually just genuinely happier because I can now deliver faster and better work. You know what a team lead of me in the past, he used to do the standup uh you know every weekly sync up or I think it was weekly and he would ask the mood like happy sad every time for every every member to measure it. >> Yeah. As part of the standup like you know what did you you but it wasn't standup maybe it was like the weekly uh check-in or something like that. But but team white everyone >> Yeah. Yeah. But that's that's a typical thing you do in a retro for example, you know. Yeah, you have safety check. >> If you have [clears throat] if you have >> if you have recorded it then you can say well since we rolled out Claude Code plans to everyone happiness has gone up. >> Yeah but >> anxiousness has gone up too. [laughter] Well, it's a difficult it's difficult to measure that too because ultimately >> I mean like personally I I I can be maybe it's the my Germanic blood but like I wake up in the morning and I'm not freaking happy dude >> and then people hey happy new year how's it going like like please >> like yeah sure >> cloud deployed all of my features that I I estimated for two weeks of work in just 3 days but at the same time the toilet on the second level is broken. It's like [laughter] >> Yeah. Oh, actually that that reminds me my particular client the the office has so many like problems. But I do think overall I think after the Christmas period, New Year period, the sentiment and the way people talk about AI has significantly changed. People are more like declaring that they don't look at their code. >> Yeah. more >> that video. I mean, you linked it to me. I can't believe he was right. >> Yes. From >> that Theo video. I to be honest, I find Theo a little bit irritating to watch for some reason. But like I did >> I always complain, but I always watch his videos. >> Yeah, I did watch his video and I thought he I thought he was absolutely spot on. I even I even uh >> You're good. >> I even tweeted. Yeah. So, I don't know. Maybe it's his hair. But anyway, the >> it's just a side comments of like but I I felt maybe I got used to them or I feel there's less. He was very much frustrated that Antropic was not giving him anything and he was very much saying OpenAI is doing a great job but maybe because it more aligns with me now. Maybe it's my little bubble and it fits in my bubble better. I was like, "Oh, he's he's hailing on Tropic now. I like to watch his videos." [laughter] But but I still feel like um perhaps I need to rewatch the video, but like um yeah, there was that that that [clears throat] whole spiel about not not looking at the code anymore. But I still feel like people I felt like he said it well maybe I'm just daydreaming at the moment but like I still there's still a lot of people who somehow throw in a weird argument that like I'm not looking at my code but at the same time I still I do look at I do look at it just to make sure it's it's producing good code like so what is it? Are you looking or you're not looking? You know, but I mean are are you like what do you mean by I'm not looking? You know, this is I feel like people are not being super honest about it. >> I think it's a it's um it's a it's a slope. How do you say it's like a a range of of looking at code? I'm looking at code. >> I'm not looking uh >> it's a range. Yes. I'm looking like at every line or I'm looking at the functions or I'm looking at the file structures or I'm looking at the you know architecture high level um you know cleanliness of of of it perhaps um but I had something else I wanted to say and now it's gone oh no that the other I thought you were going with that sentence like I'm not looking at the code but it's and then they say it's the most well- tested and validated code that I have ever delivered in my life and and a lot of people say that like I've never spend so much like effort on unit tests because and I think this is true because with the AI the unit tests are so effective now the moment the AI doesn't like write proper code in like creates a small bug it gets an immediate feedback immediately says like hey that just broke it's like oh I forgot about that and it fixes it and I think that's true I think it and we and we discussed about this for the last three podcasts and I hope I've convinced you a little bit but um but these >> I'm still a little bit dubious about even than that because in the last video if you remember we we we observed that the beads uh >> no but failing we don't talk about beats >> we don't we don't talk about beats here I mean as as a we we talk about it as a memory system not as a well tested architecture because I told you in the very early on I am very scared of updating beats because I was just looking at the change logs on the releases and it was like fixing merge conflict that corrupted all of your issue task list I was like all the I don't want to upgrade to [laughter] [gasps] you know >> I think I think DAX and open code were were extoring their tests but I haven't looked at open code GitHub okay so I think so just let me talk about >> just to summarize our last chat [snorts] >> if you want to talk about the test part because you just sent me the video of um specdriven development from AI engineer that was filmed. I looked at the date, November 22, uh, from one of the engineers that works on Kirro, the AWS specdriven development IDE, and it was very interesting because he dives into exactly where do they trust or put LLM flows and I know they because they do property based testing and if you use NodeJS, you can do that with fact check or fast check. Sorry, fast check. >> Yeah. Or hypothesis. >> Hypothesis in Python he mentions. Yes. >> So, I'm actually not very familiar with this. This is so you do you know what are you talking about then? >> So because I was porting AWS CDK to CDKTF for Terraform um or CDK for Terraform there's several um scenarios where they use property based testing in their testing harness that I also ported over. I think it was in um the way that you define schedules um for like scheduled rules or scheduled things like if you if you if you do cloudatch and you want to define a schedule and you have different constraints they do property based testing in that part of the code and I I never used it before and I was like what is all this because then I have to I have to install a fast check and I have to learn how that works and and usually I spend a lot of time trying to understand how the tool works. and how the property based testing works. But in this case, the moment I installed it and I copied over the to the the tests and then I could verify that my uh Terraform because it wasn't really testing Terraform or cloud formation. It was really testing the functionality of defining a scheduled rule. So fast check really cool. It was really nice to see it run and I'm very excited. Um >> okay. And so the fast check is a snap uh sorry >> is a NodeJS property based testing framework. Okay, >> I need to Okay, I want to see these tests run in open code and I'm just struggling to just find out how they release. Okay, how like I hate it when there's no make file. How do you do that? So one of the things that they highlighted uh in the in the specdriven development setup of Kirao is they use procedural or like algorithm um deterministic logic um to set up the testing and the validation. So they force you early on to define ears formatted tests. So uh sorry um user requirements and user stories >> like cucumber type stuff or what are you talking Yeah, behavior driven kind of like as a user when I do this I expect this or something like that. So that they are very strict in the in the way that they want want the LLM to generate these and they do a lot of like structural tests against those because they evaluate that and they generate property based testing out of those markdown. So that's super interesting, right? Because now you're no longer asking an LLM to transcribe the user story into some type of testing. you're actually fundamentally validating um whatever the user and the LLM has come to after you know brainstorming and I think LLMs are great at that um help us you know expand our horizon and think and um and then really structurally validate it and generate proper tests about it u which you can trust so that was a very nice uh aspect on it and then the second part that he highlighted was um in >> I got to see this I mean do can you show me these Sorry, my brain is not really working. So, do you have maybe you can steal the screen share and and show me these property based tests? Do you have anything to add? >> I didn't I don't he didn't show them in the video. Um I would need to install Kira and figure out where >> Oh my god. So, we we're talking about it. We don't need we >> no because we are asking the question um how can you trust the test that the LLM is writing and the way that K is solving it is by you know putting a lot of formal mechanisms around it and that's very exciting >> this before I've never seen this before >> yeah yeah blacksmith runs on um depot.deb I'm actually investigating that because I'm forking the Terraform CDK and the biggest concern is the cost of the runners because some of those uh TypeScript um programs require eight CPUs and 32 GB of memory when they are running. So I need to find a good uh GitHub hosted runner alternative. >> Okay. Do you use bun by the way? Cuz I noticed Bun got bored. >> Everyone's using bun now. Yeah, Antropic acquired it. But like I'm struggling to run this test here. >> This is bun turbo type check. >> Turbo is turbo repo. >> Type check is just a package of JSON. So turbo is a way to run uh tasks across your monor repo and and with a cache. So it will look at the input sha and then if if the input sh didn't change it will just use the cached result so it can speed up your task. So turbo is just going to run type check as a package of JSON script. >> Okay. Type check is just going to run TSC and usually type check is just running the TypeScript compiler. It's not really >> I can't work out how to run uh open code tests. >> Yeah. >> Okay. Okay. If I can't work out maybe let's have a look at the the GitHub at least. Um Oh, sorry. I'm not sharing the screen here. I'm just I'm I'm just >> There are a lot of good things from the video you shared me. >> Oh, cool cool. But I need I'm a I'm a visual learner. I need to see the bloody test. I need I need to see the test. I need to see it running. I need to see the test failing. I need to build >> trust with the the test has value. You know what I mean? It's like [snorts] >> this is a journey. I maybe it's not worth talking about now. I I kind of want to turn the conversation back to um Oh yeah. Like one thing that was very good that >> Yeah. from >> from Theo's video is the Jven's paradox which was actually mentioned by Simon Willis >> and I'm also a huge fan of Simon Willis. >> Uh if you don't subscribe to him then you're missing out. I really love his uh his tool to um share uh so so this post I wrote here is like I I use his UVX his claude code transcripts into >> there's another tool I wanted to show you as well but >> to pick out the session and then we sort of share it as a team. So I think this is like >> this is like level one. >> Yeah, this is very very level one. But but but that's where we are at my workplace. It's like it's like you know the uh the different stakeholders and the people are are basically in a browser copy and pasting between that and VS Code because the VS Code organization is so locked down. Uh anyway, so yeah paradox. Uh yeah, I'm trying to get people I mean this is a new argument for me and I I I'm I'm I'm I believe it. I'm I'm behind it. It's like if things are getting easier to do, we're going to do more of it, not less of it. I mean, I'm I'm I mean, I think it's all public stuff. I have a I have a guy that I I think I used to work with him you um he was sort of posting some like doomer stuff like it's like when AGA comes along then prime real estate is going to collapse and then I then I asked him like why do you think that and he says like with with heavy uh AGI job displacement that uh people will become on lower incomes and then people just won't be able to to meet rent I think is what he was trying to say and then there'll be like a a spiral down where everything sort of collapses and then I do mention Jven's paradox in that in that Twitter thread and then he comes back by saying that's reason by analogy. Analogy will break at the limit. It will hold for some time. um as long as we're not bottlenecked by energy supply. Uh you know, oh god. Um but but with the cost of intelligence labor approaching zero, >> there'd be very few demands that supply cannot meet. Lower incomes coupled with lower prices would would would produce a deflationary cycle like unlike any we've ever seen. Um and then then he quotes uh this some I don't know where he gets it exactly. We're going to hell before everyone goes to heaven. I mean I get that but at the same time I don't want to get behind it. I I don't want to think that's going to happen. I I mean his def his his com his combativeness of like that Jven's paradox is going to somehow break at the limit. I h I don't think I I don't think so. I mean I I don't want to think so but um yeah regarding the the the the fact that originally I always feel that we are heavily subsidized they are spending so much money and we are spend we are actually using a lot more quote than our quotas unless you're using API keys which is dumb. Yeah, I mean I I think things could get a bit more restrictive in future. Like we're definitely since since we're the astronauts of of the AI, we we're getting a free ride in some ways. >> Yeah. Which to me feels like it's going to become this elitists. Um only few people can afford I think it's already a little bit the case, right? There's only a few people that are spending $600 a month on three cloud max plans and those guys are like, you know, outputting. so much um that it's is if like >> yeah like Steve Jger is like elevated. >> Yeah, he was he was he was very elevated before with this stuff. >> Yeah, I wanted to share something that I said I would I would share a couple of tools that are really cool. Um I want to share my screen. Yeah. >> Well, actually I'm not too sure about this. >> Can you see my screen? >> Yep. All right. So, you see here at the top I have this little little bar that shows like my codecs. I've not used much right now. And I'm >> Is that Is that a Is that an official Codeex thing or is it some >> No, it has Cloud here. Oh, so stupid. 5 seconds ago, Cloud was right there. But I guess it I ran it on on uh another one. >> I'm not I'm not a huge fan of menu bars. I always feel like they're just sucking power. But [snorts] you find this one useful. >> Well, what did you >> for me from what I do right now? Because you kept asking how do you keep track of the context and how do you keep track of the usage? But I I don't care so much about context because you get a warning about 30% in. So I'm like, okay, that's when I kill the session. >> I probably have to cut this out. You I think you you accidentally shared someone's name there. >> Where? >> On the bottom left is a this >> Oh, this one. >> Yeah. No, he it's a book for about um AI changes tools, but you still owe the craft. I like that that uh page. >> Okay. I'm just a bit nervous about having to edit this video. >> It's not a CV or anything. It's a book. >> Oh, okay. >> Um so I think you'll be happy book. >> Oh, that's the author, I guess. Okay. >> Yeah, he's the author and he works at shopback in Singapore and he's kind of as an evangelist talking a lot about the tools because he's introducing them. I think he's bit of an engineer. >> So, is he is he bumping up AI? Ah, well, this is the kind of >> I don't know how how how we should reason about it. Yeah. But the fact is get back to >> maybe share that with me if it's if it's good cuz I need the reason it's open on my on my laptop is I still haven't read it. >> All right. Right. But but just just just to summarize where I am, Vincent, I need to >> to basically um become an AI mascot in my company without >> Yeah, these are definitely the recovery codes you should cut out. >> [laughter] >> Thanks. Uh, okay. Hold it. Do I I need to really cut that out. Okay, let me just make a so I can find it. Cut. >> Your name escort. You can start from there. >> Okay. What were you showing me? What were you showing me? You showing me >> Hold on. >> Oh, you were showing me this like codeex. uh you're showing me this menu bar thing you're you're flexing to the the podcast listeners or pe the viewers like yeah I've got I'm burnt I'm burnt through a billion tokens in the last 24 hours stuff that you wouldn't couldn't afford or you couldn't get through your company no that's where we were we were talking about how it's like an elitist right but that made me realize I have this little bar here at the top that shows you how much codeex how cloud and then how much Gemini I've used so far. >> How much of the Earth you've destroyed? >> Yeah, that's So people are doing these Instagram >> $500. >> No, >> that's API calls for sure because I only pay $20 a month. So So just to show that if you pay API calls, it's a bit silly because they cost a lot of money. But if you use something like um a $20 subscription, you get that in mind. But still actually these these >> but this is this is this is wrong because this is probably because I haven't spent extra usage 500 >> as a consultant people people it's it's sometimes eye watering but it's not unusual to see people being build you know thousand $1,000 US a day right for a consultant and and then and then I at the same time I get a bit you know I I shout out when I see someone that has spent $600 on Claude in a month. Of course, in real terms, AI is incredibly more efficient than a than a consultant. [snorts] Okay, maybe I have to cut that out. Yeah, I'm trying to find So, this is um cloud cuz you were showing a way to share cloud sessions. You can take a session and you can create a nice web page from it. Yeah. Using Simon Willis's tool. >> Yeah, I can link to the my blog. >> And then there's this thing called the the cloud um I don't know what it stands for like assistant session reviewer and it's Cass and it you run an indexer across your cloud session directory and it will it's using oh it's using Rust so it must be fast, right? Okay. >> Oh, come on. This is trivial stuff, but you could do this in shell. >> No, but the fact is that that like >> Okay. So, so you're basically going across all your conversations. >> Yeah. Across everything and you can type like a small little like a single word and it will find all of them. I think it's doing a fuzzy search. I'm not sure. But my indexer has an error apparent on its another problem is >> so so when it comes to these transcripts I feel like um there needs to be a tool to say like um to to rate to rate the transcripts right like like this was a really good session this this session sucked what went wrong here what went right here you know these these sort of like metadata so we have a way of sharing transcripts but we don't have a great way >> oh look at even works across open code and codex in Gemini. >> Yeah. Nice. So we don't have we don't have a nice way of of really learning as a human how to interact with AI. I feel because if we if we feed these transcripts back into AI and asking how to improve I think that I'm not too sure it might work but >> yeah but that's not >> but I'm I'm talking about upskilling teams getting getting people more to build trust with the and building shared prompts and things like that. >> Yeah. Yeah. Yeah. This is what I'm this is what I'm trying to do at work. >> So, this is one thing that they highlight in um somebody asked a question in the KO talk saying, "Can I customize the templates and the prompts and the workflow and they're like they are over here?" I don't remember the answer, but it it seemed like given the way that they are so structured, I think it's hard for you to change those things. And um I definitely feel that those templates, those workflows, like I try to use speckit to set up some type of uh ADR. So it's funny because speckit is ADR in a way kind not really exactly the same, but it's in a way you're capturing like a decision record, right? Um, so I was trying to ask specit to create a decision record and it couldn't really get it because like, oh, you're building a program. So, and I was like, yeah, no, but we're using LLMs. I was like, oh, so you're building a program that um is going to, you know, you know, you're going to call out to the open API and and I was like, no, no, no. I'm trying to do something. [laughter] So, it was really uh really difficult to get it to do kind of like meta meta spec type of things like I want you to create a library of prompts to use for my community planning sessions and things like that. It was like huh I only know how to create a program just like so you want me to create a program to run your community meetings? No, no, that's not what I want. So, it was um so yes, you need to be able to modify these workflows. Uh I tried conductor first by the way to have a feel of it. >> Oh, yeah. Oh, >> did you have a good experience like I did? >> No, actually both conductor even though in the in the product design and definition I said very clearly like what we're building here is a way to like a a type of a process for the repository. I want to build a pro. So it's a repo for keeping track of decisions and proposals and spec and so on and and um no didn't go for that. Again, probably not the right project for it. But yeah, other than that, conductor was pretty cool. I I feel it actually kept kept Gemini going like uh I thought at some point when I finished with all of the upfront planning, Gemini went in a loop and for a significant amount of time kept going. I've never seen Gemini keep going like that uh in a while. Have you had the same experience with conductor like that? It can keep keep Gemini going a bit like >> Well, it was a good experience. So I guess I I don't I mean it was a week ago if not more. So I guess I guess it was doing its thing. Yeah, I guess it was doing it thing. [snorts] Oh, okay. You know what? I need the weather is actually good right now. It's not raining. I feel like I need to wash my car and run around while there's some sunlight. So I think I'm going to end the conversation a bit early if you don't mind. I still think I still think what we talked about was valuable. I don't know if if we I guess at this point I would like to engage anyone who's listening to please comment below and uh chime in with your thoughts around making AI uh I mean are you positive about AI and things like that or another interesting thing that Theo did he had like a little poll on his video to say that do you think Opus or do you think Claude is outputting better code than you are at this point and I think it was like 65% agree agreed that Claude is already better than them at coding, which I thought was interesting. And to be honest, I agree. Like I don't know how you feel. Me too. >> Yeah, me too. I think it's better because it's better than me. >> Uh it's more like it has a broader knowledge base like it can produce code for a broader set of things that I don't know. >> Yeah. But even if I ask it to produce Go code or or code or languages I know quite well, I I don't I think it's pretty okay. >> No, but I liked Steve's take saying that it's way better at Go than at Typescript. I always felt TypeScript is perfect because it gives you the structure and the schemas, but Steve says types is TypeScript is way too you can do way too fancy things with types and TypeScript and Go is just really down to the you know straightforward. So so AI is better with Go. There there's some arguments made I think oh I can't I can't remember his name now uh the educator guy he was who makes some typescript courses he was arguing that types and AI go well together because you want those that that that test >> I always feel like that but Steve Jger was saying that that he he he rewrote gas >> typescript was was sucks in that regard >> he says it's it's a token waste it's it's generating a ton of tokens that don't really do much and uh and with >> code is so much cleaner and direct although that I like what I like with go is this very um you know write everything twice mentality of go which is like don't over generalize no generics for ages don't like just rewrite the code rewrite the functions >> I still I still don't even quite understand generics I mean I don't use them so maybe I'm missing out >> I I did originally when I was writing code golang but when I asked the llam to write I I go like yeah don't like I don't mind it being duplicated in several places cuz I'm like >> Yeah, exactly. It's cheap. Yeah. To do all the different uh >> you need to maintain it. It's going to figure out >> switches and all that sort of stuff. Yeah. Yeah. I get it. I mean, go it. I mean, I'm feeling pretty pretty uh # uh what do you call it? Inspired # grateful. because I know Go better than most people and uh and I know AI better than most people in some ways and everything everything's coming together really nicely for me. I feel and I'm touching wood here. is like everything's going well except maybe my pay packet. But uh you also did one final uh because the whole tailwind laying off 75% of staff and then people saying well well if your business model was like what was their business model right how that they got disrupted. I don't know if you know about this whole Tailwind scenario. >> Oh yeah I didn't I did notice that was that was interesting. Uh I think I didn't read the whole post and he also made a podcast himself, right? Where he's walking his >> I read the post and I listened to the whole podcast. >> Oh, really? So you probably have more details, but but my my uh what do you call it? Headline reading interpretation was that was that um yeah the LLM was was was taking their business away because uh you didn't need to use Tailwind and its abstractions I suppose to do to work with CSS. You can just get Claude to do it. >> Yeah. >> And then it boiled down also to the website. So I think there's this llm.ext and they removed it because they don't want to >> the thing that triggered it is somebody went I'm going to here's a pull request to add the llm.txt to your docs. >> And he was I'm sorry but I have to decline this. I just had to lay out 75% because the docs are our only funnel and nobody's reading them because the LLMs are just generating code. Nobody's reading our docs anymore. And it's how we we funnel our business. We've got 80% down uh in revenue. Um for you know was it 80% down in traffic through through uh those funnels. So he's measuring those different you know ways that they are driving business right and and a lot of discussions >> stack overflow problem in a way but whatever it's different. Yeah. Uh so a lot of people were saying like from Tailwind they realized as a as a an expert um I have a network and people come to me to and and that's how I get jobs but that all falls away with AI. Uh and and then some some of the uh um he says I've been sponsoring Tailwind for a long time. Love what they do. I learned a lot from them. So in his video was more like highlighting all the good work that they've done and he shared like there's this even if you are using an LLM reading their design um you know guides like they have they written a book on how to do the UX which is not just about the CSS but really about like how to build uh UX. He says that's a really good book then you should really read it and he's also encouraging people to sponsor them. >> Yeah. >> Yeah. But that I mean I'm sponsoring is always not a solution right? You need a viable business business model. If their business model is upselling through their docs, yeah, they're in problem. I mean, and interestingly, Charlie from the Astro Project mentioned that that they have LLM's text, but they also have a similar they're they're in the similar boat >> because they need to upsell their other um >> they're hosting, I guess. Oh, they have some um what is there? I can't even remember what they're business. [snorts] >> There's a there's a there's a product that I think they were working on. Oh, this Pyx thing. >> I think that's >> going to be uh a paid for thing. >> That's like Dino then. Dino and Bun like reinventing the packaging. Well, not Bun, but Dino reinventing the packaging. So, so you get the point that everything could just Yeah, they're in the similar boat. Okay. Anyway, I'm I'm interrupting I'm interrupting you. >> No, the other thing that Tio said is as a as a business, if you let LLM generate your code, go check the dependencies and see if you can go sponsor them like if but but to me it seems like the real solution is is ads, right? This is how Warp they say like we're giving you Opus 4 access for a fraction of the cost because they put >> no one's going to read the docs ads. How you >> Yeah. Because ultimately inside the docs are ads. Like if you like to if you like this here's our our offering, right? So if you translate it, nobody's going to the dogs. It means that those ads no longer live with the dogs, but they must be surfaced by the LM somehow, right? And Warp's uh operating model is like we give cheap AI access in our terminal by serving you ads and that's like and and you know if OpenAI would introduce that you know the alterate is already created but it's the only way like if OpenAI can semantically >> any business that requires ads now I guess with with LLMs in the mix LLMs are filtering out the ads >> so any business model that relies on ads is basically in trouble here and it probably includes includes Google. >> Google of course. Yeah, Google. Google had a big re wakeup call. They had all hands on deck >> uh early in the in the Chad GPT saga. >> So what is what is Google's response there >> to become a leader because they have put so much money. They have their own TPUs. They build their own Gemini CLI. They they become an LLM platform. >> But like okay, go going back to Tailwind. I mean don't say sponsorship is the answer because it's not. What is Vincent? How do you think they're gonna come out of this or how how is Astral going to come out of this? >> I think if the LLMs I guess it's it's the whole problem with LLM's not giving attribution, right? If there's a way for LLMs to give better attribution and it's like I'm you know imagine that while it's generating your website it goes and by the way there's a sale on Tailwind right now. [laughter] I'm I'm using Tailwind to build your website. No. Um, >> that sounds like another weird form of advertising which sounds insane. >> It would be funny if if if you're try if you're using the MCP or you're troubleshooting the issue and you're going like, you know, it's still not looking right. There's still an overlaying bug. It just stops. I just said, go to the Tailwind team, dude. Like, [laughter] go ask them to fix to help you fix it. >> Pay them. Um, cuz I don't have the answer here. [laughter] >> Yeah. >> Oh god. I think if a company is smart and and onropic is smart because they did like the double limits usage limits for for the whole Christmas holiday that was a gigantic success. Look at how many people tried it out. They were like I got so much done and I'm 100% in on this right now. I actually upgraded to $100. >> They just they're like drug dealers, aren't they? >> Yeah. So if there's if if Antropic is listening, Boris, [laughter] >> if they're smart, what they're going to do is they're going to make >> more correct. >> Yeah. like loop detected. Um, please like based on the context it looks like like Clippy pops up. Looks like you're trying to fix a wind a tailwind issue. [laughter] Here's the tailwind guys website. >> No, I think your proposal I mean it's frankly >> ads are coming. Ads are coming. I think this is >> there's got to be a better Okay, >> I'm going to do a hot take. Hot take 2026 hot take from Vincent. Ads in your LLMs this year. Guaranteed warp's already there. So I already what? Not except for warp. Let's talk. >> Arguably there already ads like when you run claude it gives you like a tip. Try this mobile experience. Try this. >> But it's always there. It's always like use our our designer plugin for better web design. What about it goes like use the Tailwind MCP subscription only $5 a month. [laughter] Okay. Bye. Go wash your car. Let's see how many people hate comments. [snorts] >> Okay, please please please rate the podcast. Please comment below. Please tell me how we tell us how we can improve. Actually, I did get feedback from from from Swix, but the the Riverside thing looks a bit too expensive for me at least. It's like £24 a month >> to to improve the quality. I was like, I don't want to spend money. >> Everything's improving except my pay packet. Uh, I did earn I Oh, yeah. I should I should note maybe I can give you access to my YouTube studio, but I think I I've earned two pound thanks to this whole podcast series. So, thanks Vincent for coming on the show. >> That's two pound off for a family of four. >> Two pound. I I'll I think I'll can buy you a coffee when I see you in Vietnam or something. >> Yeah. Come to Vietnam. You can feast >> for exactly one meal. >> Wow. [laughter] Yeah, I that's probably the real issue that we should have talked about is how people's pay because Madu Shan who I was chatting to on X, he was right to point out that that people's income is not improving. >> Okay. >> Yeah. [clears throat] >> Okay. See you guys. >> See you. >> See you guys. Oh my god. I say you guys. Sorry. So, see you Vincent. >> Comment below. Bye. >> You guys and girls. Come on. >> Oh yeah. Sorry.