Episode 45: We are going to cfgmgmtcamp.org

Published: Thursday, Oct 1, 2026 • Duration: 73 minutes • Season 1

Download MP3 | Watch on YouTube

We’re going to cfgmgmtcamp.org! Kai Hendry and Vincent De Smet discuss their conference plans, what makes an agent harness useful, and whether AI can reliably review infrastructure changes.

Recorded in two conversations, eight days apart: 23 September and 1 October 2026. The second recording starts at 19:20.

We compare Mecatl, Pi and the Agent Client Protocol; debate fast System One models and accountability; and explore AI-assisted retrospectives and Terraform plan reviews in Atlantis. Can an AI reviewer measure its own impact—or is it just telling us what we want to hear?

We also look at Claude Code agents and dynamic workflows, Stategraph and access to infrastructure context, and the trade-offs between Terraform automation tools. We finish with monorepos, security gates, and why better delivery depends on engineering practices as well as tools—then share our cfgmgmtcamp plans and Vincent’s Ignite news.

Chapters

Mentioned

Watch on YouTube · Browse episodes and subscribe

Will we see you at cfgmgmtcamp? What would make you trust an AI infrastructure reviewer?

summarize "https://youtu.be/aymTKw5540A" --timestamps --slides

This episode is a conversation about conference plans, agent tooling, model behaviour and practical uses of AI in infrastructure workflows. Kai Hendry and Vincent De Smet compare agent harnesses, explore automation for Terraform reviews and CI/CD, and question how much to trust an AI system’s assessment of its own usefulness. The conversations were recorded on 23 September and 1 October 2026, eight days apart.

Slide 1

00:00 Conference plans and model behaviour

They open with FOSDEM and cfgmgmtcamp plans, discussing conference culture and disagreements about AI. Vincent is preparing a CDK Terrain release and looking ahead to an AWS CDK bridge. The conversation then turns to prompting Opus: one frustrating session repeatedly produced verbose “privately” explanations before its summaries, while a fresh session on 5.5 behaved better.

Slide 2

08:28 Harnesses, agents and ACP

The hosts compare Mecatl, Pi and the AWS Strands Agents SDK. They distinguish a harness—the environment that supplies tools and runs an agent loop—from the agent operating inside it. Vincent describes extending Atlantis with a Pi-based agent to review Terraform plans against pull-request intent. The Agent Client Protocol provides another part of the picture: a way to connect interfaces to agents, rather than a substitute for the harness itself. Kiro and Zed provide examples of that separation.

Slide 3

19:20 Eight days later: fast models and accountability

The October recording begins with Google Gemini news before moving into System One models such as Jev. These return fast, typed decisions rather than general text, which can suit well-defined classification tasks. The hosts debate where that helps and where probabilities alone are insufficient: infrastructure approvals and security decisions may also need an explanation and an audit trail. An example of AI-assisted Jira ticket classification shows a more immediate benefit—making a retrospective about where the team actually spends its time.

Slide 4

29:35 Atlantis reviews and the problem of measuring impact

Atlantis can produce long Terraform plans across many parts of an infrastructure estate. An agent can inspect those plans, compare their effects with a pull request’s stated intent and post review comments. Vincent describes trying to measure whether those comments changed developers’ behaviour, only to receive inconsistent assessments from the model. Kai raises a related problem: an AI review may confidently recommend a cosmetic change as its top priority even when the change has little impact. Both examples make human judgement and a clear definition of useful work essential.

Slide 5

42:56 Dynamic workflows and agent orchestration

The hosts examine Claude Code agents and dynamic workflows, alongside Kiro’s workflow announcement. A main agent can break a goal into tasks, delegate parallel work and run repeated verification and repair steps. A screen-shared example starts with contracts between systems, then moves through implementation and integration. The structure can help a human understand the work as well as organise the agents, but large workflows consume many tokens and can be fragile: Vincent recounts accidentally suspending a session and losing a running workflow.

Slide 6

54:39 Infrastructure context, permissions and delivery practices

Stategraph’s queryable infrastructure data prompts a discussion about giving reviewers context across Terraform states. The hosts consider short-lived read tokens, separate Linux users and the boundary between an agent’s permissions and a developer’s access to secrets. Vincent then demonstrates his Terraform automation comparison site, including different collaboration platforms, access-control requirements and pricing assumptions.

From 1:04:38, the discussion turns to platform engineering and delivery speed. Installing a tool is only part of adopting it: teams also need the practices that make it useful. A monorepo security gate blocking an unrelated change leads to a debate about ownership, dependency boundaries and continuous integration. James Shore’s “Continuous Integration on a Dollar a Day” supplies an older perspective on keeping the build working.

At 1:11:21, they return to cfgmgmtcamp. Vincent’s Ignite talk has been accepted, Kai is still waiting, and they discuss meeting listeners in Ghent.

Model: openai/gpt-5-mini. Reviewed for names and timing.

Transcript (auto-generated from YouTube captions)
We are going to cfgmgmtcamp.org
Kai Hendry and Vincent De Smet
Recorded 23 September and 1 October 2026

[00:00] So it sounds like judging by your manic messages throughout the, the night that you pulled another all-nighter riding high on Opus 5.5 or whatever [bleep] the you're doing. Unfortunately, it wasn't 5.5. Um, we're gearing up towards the release of, uh, of CDK Terrain, and I'm trying to get like as much as possible, like fixed and resolved. I want this release to go out and then focus on the AWS CDK bridge. I'm super excited. I just started looking into FOSDEM and, um, what it is. Um- You don't know FOSDEM? I've never been. I heard about it. I like when I went to Osia- It's a, it's a bigger, it's a bigger event than Config Management Camp. Yeah, obviously. They say thousands of people go there, uh- Yeah ... the number of recordings and talks. Yeah. I mean, I, I love it in the sense that, um, it's, it's like a... It's a, it's like

[01:00] no one... You don't have to register, and it- Yeah ... it's a bit manic. And, uh, it's just the best conference actually. It is the best conference. The, the trouble is that there is like... There is this like feeling that things can, can be, can go the other way. Like, for example, um, there's, there's people who are like protesting against Google, protesting against AI, and so on and so forth. Oh, yeah. And like for, for example, I think Mark Shuttleworth, the, the Canonical Ubuntu South African guy, he was gonna give a talk there, and then they had this massive... They had this mad demonstration saying that, you know, no kings or whatever. And, um, he basically withdrew his, um, his, uh, talk because of it. How, how is he related to the no kings? Well, I mean, I'm just... I, I don't know what the protest

[02:00] slogan was, but it was something like, you know, no billionaires, no, no, um, yeah, no kings. Wait, who is he? You know Mark Shuttleworth, the guy that owns Canonical, the, the Thawte SSL guy. Okay. And, and like what's... Like I know I don't know him. Uh, what, what is... Why is he the... Is he a king? Is he a billionaire? He's a billionaire. He owns Ubuntu, basically. Okay. But I mean, interesting. I don't know why people are- So anyway, FOSDEM has this like interesting... Like I've never seen it before in any other conference. It's got like a, like a... It's got like a counter aspect to it. Yeah. Like an unconference type of thing or what? No, not so much like unconference. It's just like there are people who like disagree with technology at the same meeting- Oh ... about technology. Oh, shit. So I'm, I'm gonna get beat up verbally. I disagree with a lot. Yeah. I mean, I'm, I'm open to debate. I think it's wonderful

[03:00] in a way, but like of course, if someone, if, if someone decides to firebomb you or something, that's a bit much. Bit much. Yeah. It would be interesting. Um, now I hope that people, that somebody that disagrees with this podcast doesn't decide to go. I'll tell him. Yeah. Well, yeah. Like- Yeah ... I think there's, there's a couple occasions people have recognized me and, uh, and of course I've sort of reinvented myself as a- You're polarizing ... I've, I've reinvented myself as an AI advocate somewhat, and they might just have beef with me and I'll be, "Okay, cool." Oh, so that's why you're concerned about like their stance towards AI. Well, well, well, because I'm, I'm, I'm talking to you here, and I do often... I was gonna say nothing in the projects that I work on, um, you know, has a very strong AI stance. But if you're saying... I don't think people would

[04:00] recognize me anyway. Oh, great. I'll, I'll, I'll make su- I'll make sure I look different. I'll fit, I'll fit in. I stop, I stop shaving from now until February. Yeah, that's a good way to already get like targeted. Well- Saying like I'll stop shaving from f- from September till February, and I'll fit in. I like to say that I'm not scared of anything really, but of course, if something ever happens to me, I'm probably gonna be traumatized by it. It is Brussels after all. Brussels is bad. Yeah. It is a bit... Brussels is a bit of a weird place, I must say. Interesting. Um, yeah, no, I, I, I haven't been... It pro- Brussels is probably the one place I would avoid the most. Um- As a Belgian. As a Belgian, yeah. Um, I never really liked to go there as a, as a, as a kid even. Uh, so,

[05:00] so, so the dev rooms are... So dev rooms are proposed until s- the, until October, then CFPs for those dev rooms open. And I guess if you, if you volunteer or propose a dev room, then you're the manager of that dev room. Like you must have at least another person. Yeah. And you must, you must manage the program based on the CFP. I mean, I, I can maybe volunteer to help you out a bit. I don't seem to have fi- I don't, I was- I don't seem to have 5.5 on my, on my system. Huh. Oh, you mean Anthropic? Shadow banned. I just relaunched one Claude instance and, um, because it was running. It, I don't know. There's a, there's this one, uh, uh, advice in the prompting guide from Anthropic around the Family 5 model that you have to ask it, um, to explain what it's gonna do before it, it does the thing, and also tell it that, you know, the user doesn't see the, the, the shell command. So if there's anything in the shell output you want the user to see- Mm ... then you must, um, you know, summarize and repeat it for the

[06:00] user. Mm. And then things like that. And I did those prompts in one session, and it resulted in- Um, Opus 5 for the last two days, every time I asked something, every turn it would say privately, "I need this and this to verify this," and then it would bullet point list everything it's gonna do, and then it would follow up with like a, a summary of what it did-- what it saw in the last turn. Um, like, um, so far the command output has shown me this and that. And it was infuriating to read because I always have to like skip past the privately and I would jump down to whatever summary there is and then realize, oh no, I have no idea what he's talking about. Go back to the privately part, which is so annoying. Like, I don't need anything. He's, he's saying... I don't know. I should-- maybe I should show the- Yeah. There- ... the transcript. There's a lot of news. Like you, you, you're probably referring to these prompt engineering best practices. I mean, like every model has their own one practically. No, like the family of 5, like Opus 5, Sonnet 5,

[07:00] F- Fable 5 and now 5.5, they have a very specific... Like they're way more goal-oriented. They're just going ahead and doing things. Uh, people are complaining on LinkedIn like, "I asked a question and Opus just went off a, you know, a, a series of commands to, to do the thing when I was just only asking a question." Well, that's the way they've been trained. They've been trained to, you know, act on, uh, information and, and, and, and just go ahead and do things. And then I, I bet you if, if they change that behavior, you will have a whole bunch of other people say like, "I'm asking something and then it never does anything and I have to prompt it 100 times to- Oh, yeah ... to make it do the thing, right? And then at least- You're talking about user-facing . Yeah. So, so I'll show you the, the thing w- I'm talking about, about the, the privately. I, I'll just read it out. It says like, "Privately colon nothing further is needed, both at is landed and the remaining work is yours. No tool calls." And then boom, it replies with the whole thing. So it's-- it was very annoying. But, but what I did was I compacted

[08:00] and then I relaunched with 5.5 and then it, uh, behaves beautifully now. Nice. Doesn't do the privately anymore. I mean, Opus 5 needed help, man. It needed some effing help. Yeah, a lot of people were very, very upset about it, yeah. I was-- I'm, I was sick of it, honestly. I'm, uh, I'm, I'm like cancel Anthropic type s-situation at work. Um- I, I was, I was stuck with it. So anyway, there's a lot more to discuss aside from FOSDEM and Opus 5 point... Okay. So, um, there's this one thing from the... I think Joe Beda or was it, uh, Craig McLuckie, one of the Kubernetes guys and they, uh, you know, the original creators of Kubernetes at Google, um, they created an open source cloud native harness to run agents, um, on, on Kubernetes as far as I think-- I, I can tell. Like get your harness off the desktop, so it's called Mecatl.dev. Um- Mecatl

[09:00] dot... What a odd name. It's funny that they actually put the pronunciation in the name now. Like everyone's complaining about kubectl versus kube ctl, uh, versus kube... What's the other one? Kube control. So now they put Mecatl and they actually put- Oh, how do you spell "meh"? Sorry, I didn't... I'm gonna get this. Um- M-E- I'll just... It's M-E-C-A-T-L. I think it's a reference to cattle- Oh ... because your agents are like cattle. Oh, it's from Stacklok. I forgot who's behind it. Just send me the link. I've just typed the URL tw- five times wrong. I sent it on WhatsApp. So this-- I'm, I'm very into harnesses right now. I'm very interested, um, playing around with some harnesses. Um, AWS Strands agent SDK has also announced a harness now, so I'm excited to try that one out too for like running agents really. Cute dog.

[10:00] Cute mascot. Well- That's how you get 'em. I'm very happy with... Like f-for me like alternating between, uh, Claude and Codex within Orca is a, is a perfect experience for me. And then, of course, I use Hermes w-with... via WhatsApp and sometimes I use it through iMessage but I'm just so used to using WhatsApp, I just end up using WhatsApp. So I'm not too sure what you mean by why you're interested in harnesses. They seem commoditized to me. They... There's, there's no real difference to me. I mean, okay, there's maybe where you run it. I personally run my, my stuff locally inside my LAN on a Pi or, or one of my ThinkPads. What, what-- why are you so attracted to running on bloody Kubernetes? So you can have more of them? So you can try

[11:00] different things? Yeah. I'm, I'm just interested to understand why people create different harnesses, right? There's a reason why they, they create those. Like Pi.dev is clear. He said like this thing had become overly complex, hard to extend, um, and Pi.dev is minimal and you... and easily extendable. Like that's a very clear message- Yeah ... and people love it. People build their own on top of it. I used it. I'm, I'm still not sure if what I did was the right thing to do, um, but I used Pi.dev to, to do like I wanted to review, um, Terraform plans against- Mm ... uh, the actual pull request. So I extended Atlantis to, to, to, to do a review, um, and I deployed the review through Pi, um- Right ... as the harness. Yeah. Well, m-maybe I'm getting confused about the terminology here because like at work we have an-- we call them agents where we-- where they're like, you know, they have a couple

[12:00] of skills and maybe even, even an MCP connection and such and so forth, and they... and, you know, we have a, we have an agent for, for security, we have an agent for code review, we have an agent for this, that. Is that a, is that a different harness? I'm not too sure. Uh, does it really matter? I'm not too sure. But I, I, I get the idea of creating, uh, an agent personality/identity/ I don't know what you wanna call it, to do a specific thing, if that's what you're talking about. Fa- Yeah. Um, I agree with you that it's a... Oh, you, you have it open on the, on the shared screen. I didn't know you were sharing your screen. Yeah. So I think there is a bit of a confusion between what is a harness versus what is, um, what is an agent, right? Um, I agree that, you know, Claude Code,

[13:00] I would say, is a harness, right? It-it's, um, it's an environment that provides the LLM with the tools and creates an agent loop. Um, but is it an agent? I feel like an agent is something that runs more independently. You know what? I'm gonna ask ChatGPT. What is the difference between- I found myself doing that sometimes. Like, a-anytime my, my, my thoughts were not, like, clear, I'm like, "I gotta ask ChatGPT." And I feel like maybe I should just spend a little bit more time thinking about this and researching it the old way. But like, "Nah, I'm just gonna use ChatGPT." No, the thing is that I, I feel like I have a very good understanding of, like, things. I just want it to, like, help me realign, um, my thoughts and see what e- what else it throws up. Um- Yeah, like, have I got... Yeah.

[14:00] So, so this is interesting. Yeah, it's very str- Yeah. It immediate-- So this is the thing about ChatGPT and its memory, right? It immediately linked, um, my question with, with the memory of, of what I did, which is I used pi.dev, which is a harness, and deployed it into Atlantis instance to act like an agent. So is it now an agent? Why did it have to be a-- Because if you use Strands Agent SDK or if you use Claude Code Agent SDK, right, you, you will get... It's not Claude, Claude Code Agent SDK, but it is Claude Agent SDK. You get the, the core capabilities to build an agent, meaning you get tools and other things like, um, write file, read file- Mm-hmm ... um, you know- There's a lot in it ... this type of- Totally. Yeah. So- Yeah. Yes. So the strengths of SDK- The sandbox.

[15:00] Yeah. And also, like, the Strands SDK gives you, like, the ability to, you know, spin, spin off, uh, create a graph of agents, um, and do all the things. And then what they did recently is they announced the AWS Strands harness. Oh, can you, can you see my screen share? It's drawing me a diagram. Oh, it's drawing you a diagram. Did you-- This is Claude. Does it automatically use the Excalidraw or are you using some type of MCP plug? I-I've, I've got a connector here. Ah. Oh, nice. O-O-Opus fi- five point five, baby. So you found it. Oh, this is on my wife's account. It's not on my, on my en- enterprise account yet. Oh, yeah. Of course. So digital ha-harness has... I'm actually struggling to see the diff- understand this. It's not as clear as I'd like it to be. It's also interesting

[16:00] because this looks very fam- similar to a picture that, like, the Kiro people drew from AWS when they explained the refactoring they did. When they, when they did kiro.dev, originally they, they had the Kiro CLI, which is like a terminal interface harness. Then they have the Kiro IDE web interface for doing speculative development. And then they have the... Sorry, not web interface, but an, an app. And then they have the Kiro Web, which was like three different services and three different ways of, of, um, interacting with the agent and the agent loop and, and the tools that it has. And then they rewrote the whole thing, building the Kiro, I think, what they call also the harness or the core capability, um, of Kiro, and then building everything, every interface on top of it using ACP, the Agent Client Protocol. Mm. So then they built the Kiro, um, I would say harness part, and then they built an IDE, like to re-rebuild the IDE experience, the SD development experience,

[17:00] uh, using, uh, ACP to, to have basically your Monaco, um, VS Code-like interface, and then ACP drives the agent connection. And then you have Kiro Crew, which is you have your Discord-like or chat-like interface, uh, like your ChatGPT-type interface, and then behind the scene you have the Kiro AC- with ACP, it drives the Kiro harness. And then you have the Kiro CLI, um, to drive your TUI, your terminal user interface, and, and then behind the scene it drives your, um, Kiro harness with, with ACP. So I think Mecatl is maybe trying-- kinda doing the same thing. Like, what's the difference between Mecatl? Ask that in your chat. What's the difference between Mecatl and what it does versus Agent Client Protocol if you compare it to, like, initiatives like, um, like Kiro? And another famous promoter, I think, from Agent Client Protocol is Zed, right? The, the IDE that have heavily pivoted towards agentic IDE

[18:00] using Agent Client. So what's the difference between a-agents, uh- ACP and, and, uh, and how it applies to, to Mecatl. That's an interesting... Uh, there's definitely a lot more there to explore. Uh, unfortunately- Well, the... I'm sorry to-- I'm, I'm thinking of agents as like a, like a Docker sandbox and you-- and like you're, like you were just saying, you have many different UIs on that Docker sandbox. You could be Connecting it to your local desktop. You could be using it from the web. You can be using it from, I don't know, Kubernetes. Uh, anyway, that's just the way I was... I don't know what happened there. Yeah. I think that's a good, that's a good, um, you know, idea on, on understanding. It seems like you have, uh, opened the, the, the thing in, like, the artifact. You have to close the, the Excalidraw at the top right. I think it's a, a good way to think about agent-client protocol. Oh yeah, it d- it does men-mention Zed right there. Um. Mecatl is a piece of software. ACP is a protocol where the agent runs.

[19:00] Yeah, yeah. Yeah, durability. Yeah. Okay. So they- Exactly ... like what they're saying... Yeah. So basically, Mecatl could be built on top of ACP. It just, it's just a protocol- Yeah, exactly ... um, separating the front end from the back end. But the, but the protocol connects it all up. Yeah, I think this is all, this is all coming together. So today is the first of October. Google just announced, um, Argon, is it? A new Gemini Argon, which beats both Astra and Opus and, um- Really? Yep. They announced it, like, eleven hours ago. Get with the news. Jeez. Uh- How can you not be aware? I, I was watching this, like, Sky News presentation about all this AI investment and this impending bubble bursting, and it was pointing out that, like, Google was, like, putting down the most amount of cash down. And I was thinking to myself, like, if they're putting the most amount of cash down, the results don't really speak for it,

[20:00] do they? They're horrible at delivery. Like l- my- Gemini's terrible ... I had a Google, I had a Google Cloud with Vertex subscription to access Gemini. I would expect... I think it was even more expensive than, you know, their promotional bundles on, on their, like, subs- like, more user-focused, uh, plans. Um, I didn't get access for, like, months, and I canceled my GCP because Gemini was just all so horrible. They were unable to, like, roll it out smoothly across all of their platforms compared to Anthropic, which announces, and it's available the same day, you know? Yeah. So is there a pelican? How can we judge this thing? Pelican. Pelican. No pelicans yet because nobody has access to it, probably. It's like an announcement. Oh, man. I mean, in my opinion, they should just download this. If s- So they announced

[21:00] it... Oh, for the love of God. No, th- this is, that, that, that I find really irritating. Look at the benchmark. There's a picture of the benchmark. I... That's really irritating. No pelican benchmark. Ooh. That boils my blood. Well, then there's, there was Jev, which I'm, which I didn't get, uh, I w- I was a bit slow to move on, so I can't seem to get... I'm on the wait list for, like, weeks now, so I guess they're just, it, they're, they're doing some sort of exclusive- To have you. I mean- ... exclusivity play ... there's a ton of what they call system one models that are not doing text generation, but taking in the full request and sending out a single response in a typed format, which is very good, like, sub-millisecond, and, and the price is much cheaper. Um- Yeah. I no- I noticed something- There's tons of those now ... I was posting something on... about using Qwen to do it, and, uh- Yeah. So there's people that have taken OpenWeights models and, and retrained them to do the same thing and- Yeah ... posting similar results. There's Laya

[22:00] that says, "I built this a year ago." One thing that hasn't clicked for me is I, I don't, I don't actually understand how to, to use Jev in anything that I do. I don't need to make quick decisions about things. I mean, have you implemented it in something? Like, what I expect Anthropic to do is to replace Claude Code’s, um, auto mode, like the classifier, the security classifier, with something like that. I try to, uh, look at using Jev or Laya in my benchmark suite- But will, will that work- ... to, to make judgment calls on, on the output of the agents. Um, but- To, to play, to play devil's advocate there, Vincent, the auto mode has a quite a bit of like, you know, words, tracing, and then you can see why it made a decision. Surely, if you, if you cut down auto mode to, like, a yes or no, you will... How, how

[23:00] will you get the, the why? Oh, Jev is providing on each answer a probability, right? Yeah. It's, it's, it's not generating text, but it's giving you yes ninety percent, no- Yeah, but that's- ... hopefully totals to one hundred percent. ... but that's, that's harder to rationalize. That's harder to rationalize over, you know, I thought this was going to... I shouldn't be executing something like this because of the Claude MD file, and hence I will not permit this. Yeah. It depends on what you're trying to do. Like, obviously, I think a system one model, um, is way more evalu- you can evaluate, uh, it way better if it has a strict set of answers and, and a probability. I think you can-- It's way more reliable to, to, to, to depend on. Like, if it's... It's, it's very niche use case, right? It's, it's a, it's a very well-defined problem. There's only a certain amount of answers, like a set of answers. I, I disagree with the reliability. Like, to, to have something reliable, you also need

[24:00] to have that- Well, what do you mean? I, I was talking about the use case. You need logs. No, no, no, no. You, you interrupted me in the middle of what I was explaining what it's for, and you said you need logs, and you haven't even listened to what I was saying what it's for. Well, I was just retorting to the whole thing about reliability. You need accountability, reliability, and so, so, so forth. Okay. Depends on your use case, right? Which I'm trying to explain the use case. Right? I was saying that people like Jev so much because of the use cases they use it in are there's a very f-very well-defined set of answers. Um, they only want one of those answers. It needs to be well-structured. It needs to respect that structure. It needs to be cheap. It needs to be fast, right? For those use cases, that's where they want the System One model. They are doing samples and demos of how powerful it is with like, "Oh, look, it can play Tetris," because it gets in the input of the game, and it responds with the next move within seconds, right? Um, with not

[25:00] within seconds, within milliseconds. Yeah. So it's so fast that it can even play a game. So i-it is reliable because you have a, a system that needs to run and, and you know, an LLM takes time to respond. It generates an answer, and then sometimes it doesn't even-- isn't even valid JSON. You need to add like a whole bunch of gates and validation around it to make sure that it's actually valid JSON and so on. So you can't trust the, the response of the LLM, which is a text generator, right? So a System One model use cases seem very well-defined. Now, to come back to the argument that you're making, I said it could work for an auto mode. Uh, you said absolutely not because you need audit, you need, you need, uh, this and that. Okay, fair. Not the right use case. Yeah. Okay? I was just saying I'm waiting for Anthropic to, like, employ it where it's, where it's appropriate. I thought maybe auto mode, you said no. And I had a very similar experience when I suggested to Opus, how can we use, like, the System One model? Because I have a lot of problems with validating the output of the agents. I'm running OPA, Reg Val-- um, Reg, Rego, um, to, to

[26:00] identify if the agent is building an infrastructure according to certain grades. Um, it calls it oracles, and they're at, like, static all the way up to live, which means the infrastructure has to be deployed, which means, uh, it mutates the account, which means you can't run multiple tests at the same account unless you separate them. So, so I have these static validation of the, of the, you know, agent's, uh, results, and I'm like, "Couldn't I do this better- Mm-hmm. -with a System One model?" And Opus was like, "Absolutely not. You're gonna have all of these problems. It's not gonna make any sense." Um, so again, probably not the right use case. Well, it seems, it seems very ni-- okay, well, my takeaway is that Jev and this type model responses seem very, very niche compared to the general use of AI that we're all accustomed to at this point. So... I think it makes sense if you need to run agents in production, which is, like, if you need-- if you have a system that needs to do some type

[27:00] of intent, like needs to take in a lot of context. Yeah. It's very hard to do like a pure algorithmic, um, you know, to va-evaluate it. It has to have some type of intent or s-sentiment judgment. Yeah. And it needs to make a, a quick call. Like, like a classifier or something like that. Yeah, like email classifiers. There's a lot of people coming forward say, "I, I run Jev on my Gmail to classify thousands of emails, and within, like, a second, it classified one thousand emails- Yeah. -and my email inbox is now handled." We're, we're actually-- funny you should mention that, but, like, recently I was impressed. I'm always impressed when... I know this sounds very disparaging, but, like, you know, like, you work in a company and there's a, there's a product owner or there's a scrum master, and he's obviously non-- or he or she is obviously non-technical. And then they do something with, uh, Claude. I'm like, "Well done. You did something." Um, so, uh, I mean,

[28:00] I'm being disparaging, but really it's cool. You're being mean. It's, it's-- I'm being, I'm being mean. But, but seriously, you know, the, uh, I can't remember what is-- what this person's role is exactly. This person m-managed to get, you know, the script runner JIRA tickets and put it into Claude, create an artifact that, that was interactive in the sense that you could, you could change the parameters a little bit, you know, by tag or s-some filter, and then, and then it would classify the, the JIRA tickets. Much like I think you suggested, uh, to me some time ago, is like it was going back in history to sort of like work out, you know, the areas where the team has been spending time. Like the team has been spend-- you know, twenty percent of tickets are PR reviews, twenty percent of tickets are responding to, uh, this team and their asks. Twenty percent of the time is done, I don't know, doing, um, some other thing. And, uh, I thought that was really interesting and,

[29:00] uh, like this person just did it without me suggesting it. And he- Well, there you go. And the person did a great job. And I was like, I was like, I was like, "Yeah, this is the power of AI." Yes. Yeah, this is great. This is-- th-th-this is the power of AI. So the retrospective was a lot better. He brought it into a retrospective. So like, like, "Oh, what should we work on? What can we improve on?" And like, it was the first time, instead of like a retrospective being like, you know, "How do you feel today?" It was like we had the numbers there, and the, the numbers-- it wasn't cold, hard numbers. It was classified numbers, so it was, it was really actionable. They, they've gotten better, um, but it's also scary to... I, I tell you what my experience is, right? I, I, I might have mentioned this last time, but I-- So I basically have Atlantis, which is a Terraform automation, uh, and collaboration tool that, that runs whenever somebody submits a pull request to this Terraform repository.

[30:00] If they make-- or to their product repository, they make a Terraform change, Atlantis will run Terraform Plan over it to evaluate the impact of that change against the live infrastructure. And then that plan can be very lengthy, right? Terraform plans can have a lot of information in there. Uh, on top of that, if they make a change to a shared module, that triggers, um, a plan across all of the instances that module is used, um, if they are, you know, trunk-based. So if somebody makes a change, you immediately see the impact across the whole infrastructure, um, if that change is gonna cause problems. Does it-- It's been a while since I've used Atlas-- Atlantis. Does, does Atlantis-- can it run plans in parallel and things like that? Yes. Okay. Yeah, yeah. But that, that brings its own problems with it. Um, and it's also one of the things that- criteria that evaluates on the, on the website, um, you know, that, that I built to evaluate Terraform automation solutions, and that has been featured on the, uh, Terraform weekly newsletter as well last week, which I was like, "Huh?" I built it for my own. Anyway, the,

[31:00] the thing was I was, I was trying to say is, and where I wanted to, to come back to the, the, you know, using, uh, Sonnet or Opus to go through a bunch of, uh, re-records and make an analysis. Um, basically, the plans that Atlantis generates, in some cases, it gets posted across, uh, in, in a pull request comment that has a limit. Um, so then it gets split, and then you have like five pull request comments. Each one have like a part of the plan. You need to go through it. Um, it's, it's complicated. There's a lot of churn sometimes. And so we miss things, things that cause problems, that we looked at the, the, the Terraform plan, but because it was so much, we, we, we missed the detail. Yeah. So, um- Yeah. Of course, the obvious thing is, yes, um, I want to put a little... You know, Atlantis has a really cool feature. Atlantis is, is, is basically an optimized CI/CD, uh, for, for Terraform, right? It understands- Yeah ...the concept of state. I was, I was gonna say that I'm in this b- a, a very large Terraform

[32:00] estate, and we use GitHub Actions, and it's a freaking nightmare. But- Yeah. So the, the difference between Atlantis and- Atlantis is way better ...and, the difference between Atlantis and the GitHub Actions is that Atlantis is aware of like keep, you know, a lot of the details that- Mm ...you know, a generic CI/CD doesn't have. Yeah. So what I'm trying to get to is that Atlantis is extensible, very much like GitHub workflows. You can extend the, the, the built-in workflow, the default workflow. The default workflow is plan, apply if approved. Um, sorry, plan, do a Conftest, like an OPA evaluation, apply if approved, and merge and auto merge the, the PR and, and discard the logs, uh, when it's done. You can extend that default one, and so you can add pre-workflow, you can do a bunch of things or post-workflow. So I extended the post-workflow. After the plan has ran, um, it will then look at every plan that's on disk. Out there. Yeah. And then it will pass all of that into an AI, um, agent.

[33:00] Um, I, I'm actually running a little agent, like a harness. So it's actually, it has tool calls. So it gets a little brief of like, "Hey, here's the, the, the, the pull request body. Here's, uh, all of the plan files on disk that with a little summary." And, and then it can go ahead and, and go and look and, and find out details about, about this plan and identify if there's any concern. So it compares the pull request intent against what's actually in the plan. And getting to the point - No, no, I'm with you ...but the point was, it was running for two or three days, and I wanted to know- What? Okay. The... It, it was running for two or... No, no, like it had been live for two or three days. Okay. What I'm saying, not is it keep running for two days. I, I put a, a, a two, uh, a turn limit of like fifty turns. So there's a limit on how many turns it can take before it posts a comment. So it's not gonna run for two days. It was live for two days. And it... I asked Sonnet, "Go through all of the pull requests over the last

[34:00] two days. Look at like how is it like... Is there any, any, um, sentiment that we can derive? Is there any impact?" Mm. And, and Sonnet was very, very eager. It's like the f- after two days, it was like, "Yes, there's this PR. Um, the, the AI comment reviewer was posted, and then we see, uh, right after another commit was pushed addressing some of the, uh, the findings." So that user saw the, saw the AI review and acted upon it. Mm. And the, and, and it has impact. Um, so, so I was very happy, right? I was, I was like, "Okay, cool. I'm gonna, you know, sh-show this to the rest of the team." Um, but what I'm, what I wanna come to then is then I made a change, and then I, I deployed it, and I let it run two days, and then I asked again Sonnet. I thought, "Great, you know, give me one of those beautiful reviews again," right? And it says, "No, no impact. Actually, people are completely ignoring it." And I was like, "What are you talking about? Like two days ago, you said it had a lot of impact." And I pointed it to that review, and I gave it some hints of how it did the review, and it says, "Oh,

[35:00] you're right." And then it do the review again, and it says like, "Oh, actually on this PR, people reacted and like, uh, updated the PR body after the review," and blah, blah, blah. And I was like, "How can I trust this?" You know? You see the problem and like- How are you me-measuring impact when people like upvote the AI re- AI, uh- It doesn't, it doesn't upvote. But it, it would say like, "At this time, the, the comment was posted. At this time, another com-commit was pushed." And it would say, "And the commit is addressing the findings from the comment. At this time, the PR body- Mm ...was updated." So Sonnet was doing all these queries on its own. Opus or Sonnet- Yeah ...I can't remember. But, but wait, I don't really understand this logic here because maybe, say people get better at creating commits first time and then, and then the reviews- Yeah, it doesn't matter ...become more and more- It doesn't matter. What I'm trying to say is, um, what you could, you could argue that I didn't do a very, uh, scientific approach to my analysis of, of the, of the, of the impact. And I didn't. I asked Opus, "Can you find anything? And here's a couple of PRs- Yeah ...and here's a couple

[36:00] of examples of what I consider impact." And then Opus came up with a couple of rules and ran those, identified all the pull requests where an AI comment had been posted, um, by filtering by the name and then, you know, came up with a couple of things. My- what I'm trying to get to is that AI can be very convincing at giving you a sentiment analysis over the last five days, um, but how much can you trust it? Like they- Yeah. What I was saying is you can... They have gotten a lot better at it, but I had some experience where I was like, "I don't know if I can still, you know, trust it." Well, actually, I think you, you, you're hitting on a, on an interesting point because like I, I usually clone out a project and I go like, "How can I improve this?" or something like this. And sometimes it, it finds real clangers. But obviously in some mature projects where things are pretty under control It says, "Oh, you really need to fix this re-- this problem. This is your number one priority," which is like, I don't know, a name change or something, which, which in reality has

[37:00] no impact. Like, it has zero impact. Like, I'm trying to get AI to find the most impactful changes, and it's suggesting me things that really have no impact. And, um, and, and that's one of the things that, uh, I feel, yeah, it, it becomes like a trust issue, but you have to like... You just have to... You have to, you have to set the benchmark for AI and, and realize that it's, like, just not doing the right thing or not, not having an impactful change or just wasting time, I suppose. I wanna share something else. So we were saying earlier, uh, you were saying that the output is not consistent, sometimes very good, sometimes very bad. Unrelated to what is on screen right now is a post that I saw on Reddit today, which is, uh, a guy who says exactly same prompt one week ago and today, and comparing the results. And, and today was a lot worse. And people

[38:00] are like, "Is this nerfing?" Is Anthropic, um, you know, really boosting the performance, um, of the model on, on the release day and then a little bit later to, to, to save cost is, like, nerfing it or, or maybe they call it quanti-quantis... What is it when they are quantizing it? Yeah. Anyway, um, we never know. And then people are saying maybe when you, when you select medium on the launch day, that's like a hundred effort. Yeah. But when you select medium on, on, like, two days later, uh, two weeks later- I definitely have that same- It's only fifty ... I have that same feeling that they're, they're messing around sometimes. Yeah. So, so that's where-- But I'm, I mean, I'm sti- I don't have that, like... Okay, avoid these, you know, avoid these frustrations and don't run the same task twice. Simple. Don't like this. And don't, don't look back. Just keep looking forward. Then you'll be a lot happier. Okay. But this is very interesting. Anthropic, uh, did this article, uh, I guess two days ago, right? Mm-hmm. Um,

[39:00] where they talk about how Mythos, um, you know, was running on a, on a exploitation benchmark. The, um, they have an, a benchmark to evaluate how the model can find exploits. And Mythos was scoring like, I don't know, um, there was-- They, they highlighted the score that Mythos got back in March. Wait, what, what does GLM mean again? Sorry, I should know this. GLM 5.3 is a model from Z.ai. Um, it's, it's one of the- Oh ... like- Oh, so it's benchmarking GLM- Yeah ... with, with Mythos. Okay. So the latest release of Z.ai is GLM 5.3, and it does better today the open weight model that you can run or that Z.ai offers on their inference engines, um, and it-- at a lot much cheaper cost than, than Mythos. Mythos is not even accessible. They say 5.3 today scores better than Mythos in March. Um, and, and they didn't release Mythos to the public because it was

[40:00] able to, um, bypass safeguard, like able to find exploits in- Ah ... um, in like software. They actually have... All the numbers are in here, you know. Um, what is the actual exploits that they're doing, like, uh, OSS, Fuzz and, and things like that. And, and what they're saying basically is, you know, it doesn't matter that we are protecting and not releasing Mythos to- Yeah. The other ones are catching up quickly. Yeah. Um- That's kinda scary. It is scary, but it also means that you know how sometimes when you ask Opus to do something and then it drops back because of safety guards, it's like, "Oh, you're trying to do something that triggered my safety guards, so don't be upset, but we're downgrading you - Yeah ... to a lower model." It means that I shouldn't be-- Like, if I have any, any cyber or security task, I should be using, uh, Z.ai, right? Because- Yeah. It's funny you should mention this because I hope

[41:00] this is, uh, a n- a hope this is non... I hope this is public information, but, but basically Claude has a security product that they want my client to use. Yeah. That's also what I said. Yeah. They're obviously selling a security product- Yeah ... for the premium. And now Z.ai is- And, and, and, um, you know how it is in big companies, like, "Oh, should we just turn it on?" But, like, we have a very large estate, and I'm thinking, um... And, and, and also we, we, we use Datadog a lot to basically manage, like, a single pane of view. And I think the way that the Claude security one works is that they have the, their findings in their own whatever panel. It's not integrated with Datadog. So I was, I was initially proposing that we do the triage and, uh, we use-- we basically just use raw Claude and feed into Datadog to build a b-- to build a, a picture. Like, when I say feed into

[42:00] Datadog, use Datadog MCPs to sort of mute false positives and things like this. So, so there was... What, what I was thinking of is just using MCP and Claude. But then Claude is coming across and saying, "If you just switch on this button, um, then A, we'll charge you a lot of money, and B, everything will be secure, or we find all the securities in their thing." Um, I, I don't have control-- I don't have, uh, access to their thing, but I just can't... You, you... The point I'm trying to make is that I'm just not, I'm not keen to use their, their thing really because I feel like, A, it will be expensive, and B, it will be disconnected from Datadog and the rest of our operations. But- Mm-hmm ... it'd be interesting... If anyone's using, uh, Claude Security, comment below and tell me if it's any good. Some sunken cost fallacy. It's like, has to be good. Um- For that price. Paying tons for it. Yeah.

[43:00] So there's another interesting announcement yesterday, which was, um, Kiro announced, uh, Workflows. And when I read this post, it really sounds like dynamic workflows from, um, from Claude Code. But it seems that they're doing a little bit more, um- Oh, yesterday a colleague showed me that he only uses, uh, Claude Code using agents, and I didn't even know about this. Did you... Do you use agents? I mean, every time I ask to, to do something, it spins up a workflow with like 20 agents. So do I use agents? I guess. But, and, and, and do you use it with... I think there's like a Worktree option. So basically he uses- Yes ... Claude like I use with Orca, and I was like, "Oh, I didn't even know you could do that in, in Claude." Okay, fair. Ah. I'm a bit late to the game. Well, I-- Claude Code has an enter Worktree tool. So I can tell, for example, I have, I have Fable. Sometimes

[44:00] I, I run three different things and I say, "Okay, you spin up, enter a workflow, work on this pull request, um, and, you know, run..." It runs a dynamic workflow- Okay, so you do that all in Claude. Okay. Into a Worktree. I guess- Yeah. I never- Claude Code, yes. I never figured out how to do that, so basically I ended up doing it in Orca. Oh, there's also this one thing, um, that I did use, which you can do... And I sh- I talked to you about this before. This one, right? Yeah. Yeah. That, that's, that's what he showed me. So he just- Yeah ... he just bounces between all these agents and like, holy shoot. Yeah, I disabled- I didn't know about this. I disabled it. Um, first off, what I found very frustrating when I used it, you know, months ago, it's that I don't know which work that-- like which repository they're working in. Like this, when they announced it, they have like a little icon of the, the, the little crab or whatever it is, using a lasso to pull all of the different sessions into one-

[45:00] Mm. ... dashboard, which is this dashboard. Now, the problem I have is I don't know where these sessions are running. Like- True, true ... one session is running over here. I think Orca is a lot clearer about that. Yeah. Maybe they've improved it. Uh, and then another thing happened, and there was a bug. Um, normally, if you're inside a session, when you press the right arrow, you go back, you go back out to this like overview dashboard of all of the agents that are running. And what happened is I had like a massive dynamic workflow running for an hour or so, uh, and I pressed accidentally the arrow key, and when, when it... That happened, it suspended that session and killed the workflow that was running in it. And I was like, "Goddamn it." Like, it was a, a workflow that was running for hours, and it wasn't done. I was very, very pissed. That's annoying. But I actually- But they fixed it. I'm actually struggling to understand the relationship between this thing you're showing and dynamic workflows. Is that something that, it gets spawned from an agent view? I mean, every session here, like even, even

[46:00] a session, um, that's active- This is so confusing. The terminology is not great, is it? Every session, like even this session that's here right now, I can go into another shell, and I can resume, and then I will have that s- I'll have my terminal connected to that session in two different, uh, places. So under the hood, they're all running as agents. They're... You can all, uh, suspend them, you can all put them in the background, and you- Okay. I mean, maybe, maybe it gets activated when you run the agents command, but for me, each session is an agent on its own. Uh- So what does dynamic workflow mean? It, it sounds like another breakdown of an a-agent session. No. Dynamic workflow is, um, basically it gives the main agent the ability to define a predetermined step, um, workflow. So, well, we can look at some of the dynamic workflows that are running here. Um, is it this one? This one just, this one just landed.

[47:00] So if I go to Workflows, um, I have a couple of them here. My lord. And so what you can see... Not this one, that's not interesting. Let's do this one. Okay. So this one ran for one hour and nineteen minutes. It ran 17 agents. And you can see each one is like Sonnet 5.5 and the number of tokens. And it first implemented, this is in a monorepo with like multiple systems, and these systems are talking to each other. So it implemented the contracts, like the interfaces between the systems first, and it did those two in parallel. So it did the first contracts, and then it also did a couple of other things. There's another one which is super interesting. It's probably this one. 'Cause this one, the contracts were really nice. Like it did 12A identity trust contracts, and it did 12, uh, 13A agent execution contracts, and then it updated the SDK. So just talking aloud here, it looks like dynamic workflows is more of a structured way of running your session or something like this.

[48:00] Yes, correct. Like, basically, you have a main agent like Fable or Opus 5.5 can be your main orchestration agent. Then what I usually tell it is like, "You're the main agent, your context is precious. You focus on the long-term goal, what we are trying to achieve." Mm. "You achieve that goal by breaking down into smaller like iterations." And then depending on the iteration size, it can either be a background agent that does that iterate, that one iteration, or it can be in a dynamic workflow, which will be a large, um, you know, number of things that need to be de-de-developed. And, and a dynamic workflow, the main agent can then decide what needs to be done. Like in this case, we are building identity and trust and agent execution and then, um, something else, um, as a platform, and we are defining the contract between the two systems, then we're implementing the, um, you know, the actual implementation behind the interface. Then we're integrating across those different, uh, implementations- Mm-hmm ... and then we

[49:00] go into a verify, repair, verify, repair, verify, repair. Yeah. And then normally it will do a comment linting, uh, step where it removes any comments without debate. I need to try this, but I th- I think dynamic workflows is disabled on my thing. Yeah. So dynamic workflows has been there since like May, and I talked about it several times because, uh, it got rid of half of my spec-driven development s- things. Because they're basically doing the whole breaking the problem down, breaking the plan down into tasks and phases, because that's what happening here, right? It took the plan of what do we need to build, and it broke us down into different phases, different tasks. Yeah. Each agent has a different role and, and in this case, you can see it says, "Okay, we're gonna run it, run with the Opus," and it, it, um, it verifies, and it then gives a review of like, um- Yeah. I, I, I mean, I, I see it more from a- Here are the open questions ... the perspective of a human, like as a human looking at this task that you just showed me you're working on. It has some bullet points or whatever

[50:00] to, to see how the problem was broken down and that, that alone, from a human perspective, has value. Never mind, never mind the other way around where the agents were like, uh, you know, put to work in an efficient way or something. Yeah. And then, and we also talked about like agent, like do you create agent profiles? Uh, do you say you're an expert security reviewer or not? Blah, blah, all that. Um, I don't because the dynamic workflow kind of on-demand creates a prompt that says, "Okay, you're the verifier here. Uh, this is what you're gonna get. Uh, we're in this thing, you need to focus on this." And, and so it basically, the dynamic workflow on-demand prepares an agent, selects the model, and, and then gives it a task. That's what your role is. Like you run a read-only, uh, review, um, you know, step of a dynamic workflow. And, and they have really improved the, um, capabilities here because

[51:00] these agents can, they can send messages to each other, they can do like many things. Like the main agent gets messages from the background agents. Um, I mean, dynamic workflow's been in there. Uh, I don't think a lot of people use them because they are so expensive when you, when you first start. Uh, when they first came out, people were like, "I ran one dynamic workflow and my five-hour session was gone," 'cause they spun up 20 agents and they used so many tokens. So, um, yeah, anyway, the reason we talk about dynamic workflows now is that Kiro, uh, just introduced it and, um, it has a lot of the same ideas. Again, it's a tool. Like there's a recipe. The, the main agent knows how to create those dynamic workflows. Some of the things that Kiro seems to do different is that they can update the dynamic workflow. But even that, like Claude, Claude did it here as well because he saw that the loop verify, loop verify was stuck on one thing and he paused the workflow, and then he relaunched it and then continued. Mm. So, so, so it is dynamically adjusting. Like even

[52:00] the main agents can see if the workflow is getting stuck and then like stops it and then relaunches it with different steps. Okay. Can we go on a little bit of a tangent here? 'Cause I just, I'm just thinking back to reviews and pull requests and things like this. So one problem I have, uh, it goes back to the Terraform stuff you were t- mentioning before. One problem I have at work is that we, we, we make a Terraform change and there's a bazillion plans and they're very difficult to go through and summarize. Um, a- a- and even, it gets even more complex than that. Like, like sometimes a Terraform provider up- has an update and there's some s- there's some very subtle changes that we need to be aware of when, when we, when we approve these, these changes, right? Mm-hmm. So what I'm trying to say here is that what is the optimal flow for review? Because, because we have a lot of changes in Terraform. It's like, it's not like I want to be a s- I mean, I do wanna be a superpowered human, but at the same time, I kind of want to make sure that the,

[53:00] the CI had, has a dynamic workflow itself is, is, is how I'm tying it back to here. I don't know if you've ha- have any thoughts about, about, uh, orchestration to help. Okay. So what you're talking about here is like basically orchestration of the, of the, of the plans. I- Yeah ... I'm saying that like work, like workflows are- I mean, plans is just one example, but like it, I think a plan, plans is a good example because like what happens when you do a plan? It sh- shows you what the changes are. There could be lots of changes and then you need to synthesize that to, to make an approval. Okay. So you're saying I want, I want, uh, Terraform cross-state orchestration, but I want my agents to review it. Yeah. Uh, are you t- A- and basically have all that context because I, I know just... and like, and there is like copilot, uh,

[54:00] review of course and, and all that other stuff, but usually it only looks, it only looks at, um, the code change, right? It doesn't, it might, I think it might look at the actions, I'm not too sure. But y- you know what I mean? Like the, the, um, there's a lot going on and, and I feel like I need to be intentional, I hate that word. I need to be intentional but, about building a Terraform review bot that's able to know about pr- provider changes, just know about a whole bunch of context about the organization and then make these approvals, uh, ideally on behalf of a human. I think that's why I'm very excited about Stategraph because Stategraph, they are really giving agents very interesting context for, for this type of, um, um, activities. They're able to... Because they're not constrained by like, um- You know,

[55:00] having one state file per Terraform workspace, they are basically having a very large, um, database, and then they allow the agent to do a query to, to basically identify if I make a change in my networking layer, like what are all the resources- Mm. -that actively depend on this? What will be the impact across my whole infrastructure? So when they, they... When Stategraph, when they, uh, broke down the state file and recorded each and every resource as a, as a row in a database in Postgres, um, and then they added a whole bunch of like, uh, query capabilities, so the agents can run all kind of simulations and, and it's in their docs. Some of the capabilities that they make, um, possible for, for, um, AI are actually very, very exciting. Um, in this case, they are showing that you can, you know, generate a token that's only read, uh, that can only-- that the agent can use for, for

[56:00] doing a-an analysis, um, and it's short-lived. Um, but they can do a lot more. Like, they're really building this to be more agent-friendly to- Yeah. -to be able to query the, the database and, and do these type of things, these queries. And I wonder, I guess the, the general practice for secrets in a state file is just to not allow the agent to access just that part of the state file, right? So if you run- I wonder what- If you run- I wonder what guardrails there are there. If you run any binary like OpenTofu or Terraform, the agents do not see the secrets, right? Um, they are masking sensitive values. Um, if you say that- Oh, so sorry, I didn't... Can you just repeat that? I did, um, my brain- Why would you... Why would an agent read a state file? You would not. Like, you would run a Terraform binary or an OpenTofu binary, and it would mask secrets. Right. Right. Well- It would never add into the secret or into the context of the agent. Well, uh- No, no but. I'm,

[57:00] I'm just thinking that agents are really smart, and they might just wanna grab it for some reason. They should not have access to it. Um- But the trouble is they're usually running under a human developer's context, which has access. Why does... You know what I'm trying to say here. Staging and production, in our case, uh, the, they, the humans do not have access to it. Like, the only way that a Terraform plan can run in staging and production is through a pull request, and then it runs through an Atlantis instance that has the ability, yes, it has the ability to decrypt and read a token to, to apply the change into the Terraform. Um, but the agent, I mean, the, the, the instance is isolated. No human can run it. And then when I'm in-- for my AI reviewer, it actually runs under a separate, uh, Linux user isolated by the kernel, and it, um, it doesn't even get access

[58:00] to those tokens. Like they are, um- Okay. Yeah. I think, I, I think the problem is I probably have too many per... I, I have too much permissions. I need to, I need to basically- That's what, that's what Terraform... Okay, so a lot of people use Terraform- I need to switch my role or something. They think when they use Terraform that you don't, like, that, that everything is hunky-dory 'cause you can run Terraform Plan and Terraform Apply. But they never realized the complexity that you get into when you actually need to manage Terraform across a, a large team- Yeah. -a large enterprise, and you need to have- Well, it's crazy. -role-based access control. Especially with, with the networking, 'cause if you F it up, it's a nightmare. Yeah. So that's why there's this whole, you know, um, ecosystem of, of tools that are built around, um, orchestrating infrastructure as code that are obviously not free. Um, and that's why I built this website that, um, basically compares all of these- Yeah. -solutions. I added a new feature.

[59:00] You wanna see the new feature? You don't wanna see the new feature. No? Uh, first, first, actually a lot of new features. Okay. Um, because some of these, they say like they support Slack, but they don't support, uh, Microsoft Teams. So I added a little toggle because I don't think it's fair to, to grade this, um, these... If you're not even using Microsoft Teams, you don't care, right? So if you're using Slack, that changes the score because, um, some of these have- Mm. -better support for Slack than for MS Teams. If you're using Bitbucket, a lot of them fall down 'cause they don't have good support for Bitbucket. They mostly support GitHub. So s- Who, who, who is to judge about... I guess you, you need pull requests and... Oh, God. What are you talking about? Uh, I'm just wondering. I was just thinking out loud, like what, what do you mean by good support of Bitbucket? I guess y- it's more than just the Git protocol. It's- Right on their product page, they say, "We don't support Bitbucket." Because again, Terraform automation- I didn't know that. -and collaboration solutions are often tightly coupled to your v-version control system, right? Mm. They might have a GitHub app. Right. They need to

[1:00:00] react onto webhooks. They need to react onto pull requests. Yeah. So obviously, integration within the, the version control, um, you know, platform is very important and is not equal, not equal for every one. Okay, so the next feature that I added, sometimes I don't... I, I just wanna know who's the best at like, you know, RBAC and SO, uh, single sign-on, right? I don't care about like the overall weighted score. I just wanna sort it by this one particular criteria, who is the top performer in that area. And so here it's showing env0. These are all the ones that have, uh, like the highest score for this particular criteria. Um, and so that, that was the first, you know, little itch that I wanted to scratch so you can- I don't, I don't quite understand these little chiclets. So if you, what, if you go, if you go to the zero, does it tell you what the partic... I guess you have to hover over it. Yeah. Oh. I don't like the, the width of this- Yeah. -of hover over. Yeah, like normally- You don't have to hover over it. You c- if you go

[1:01:00] down, you can see the full thing. Yeah. Um, but still that's not ideal. Hold on, hold on, hold your horses. Hold it. Hold it. Um, you can also If you really wanna say, "Okay, I've made my selection. It looks like I really wanna s- I, you know, I'm most likely, if I look at pricing, um, based on my team, the number of seats that I need, the number of- Mm-hmm ... resources that exist in my Terraform state and so on, I realize that, you know, env0 is extremely expensive. It's like $11,000 per month. So I decided that I'm not gonna use env0 because it's a top performer here, but, uh, actually Spacelift and Stategraph are the two that I wanna compare. So I can switch this over to, um, Stategraph, and now I can actually really see what matters to me, which is like this is the two that I really wanna compare, and I can see which in each criteria- Right. Right ... why are, is this one considered like better in that criteria, um- I, I mean- ... compared to the other one. I find this m- like the, the current enterprise I'm working in, we use GitHub Actions. I feel like

[1:02:00] GitHub Actions needs to have its own uh, entry here just, just, just so that I can point out- No, it should not. Absolutely not ... to my colleagues who don't know better. I think I might add a row like GitHub Actions is absolutely not. Uh, like it's at the absolute bottom at criteria zero everywhere. Yeah. I mean, this is great when you're choosing, but like, like for me in, in this enterprise, it's like how do I, how do I get them to use Atlantis a- and have... Oh, so what you're saying is I want to show them s- uh, head-to-head, um, what GitHub Actions can do and what others- Like just, just to show you what it buys you, what the migration is between the two things. Okay. Okay. Because yeah- Yeah ... I'm in this world of, of just using bare bones GitHub, and it's... I do miss this tooling around 'cause, 'cause I s- I, I see the pain that we have whenever we, we're, we're trying to rolling out a, a change across the ecosystem. Yeah. So this- Is Atlantis in h- in here? No. Yes, Atlantis is in here. It's like it's $90

[1:03:00] for my self-hosted instance because it only runs during business hours, which is also a security feature. Um, you can only- Interesting ... p- plan or apply during, uh- Interesting ... the weekdays. Yeah. I need- Um- I need, I need to fix things in this employ- in my client. I need to fix things, but I don't know if I have the, I don't, I don't know if I have the clout to, to get people to move to a better system. So I actually have Atlantis as, as a, as a construct, so it's like, um, here, but it's not, it's a private repository. Mm-hmm. But this is the actual AMI and Terraform- Okay ... module that I deploy and run. So this one costs me $90 a month. Um, I, I am also adding the AI reviewer in it. That means you're getting, um- Yeah ... you're basically getting all of the like top-notch stuff, but you can self-host it. Um, you get your, you know, you get like a ton out of it. This the, this is the, the important use

[1:04:00] case I feel because everyone's sort of bottlenecked by approvals and moving, uh, m-moving w-with, with speed. They're getting velocity, right? If, if you can focus, I'm sure you do, on what would basically unblock the pipeline, i.e. AI reviews, then I feel, uh, you know, what has... Which, which system out of all these things would you say on a, on a large estate would you get the highest velocity with? Like what has the best AI review approvals that, that are proven to, to work? Okay. So there's quite a few talks that, that been on engineering AI and so on, and they talk about AI, SDLC, and velocity and in-increasing velocity. And if you ta- if you look at those talks like Uber talking about it or others, the first thing that they say to achieve that velocity, they adopted monorepos. They spent a year rolling out trunk-based, uh,

[1:05:00] you know, promotion pipelines. Yeah. They do a, a ton- Getting pipelines maturity ... of a front platform engineering effort to then, uh, actually achieve high velocity. Mm. And, and, and I think somebody was like, "Oh, you know, AI, SDLC, dark factories, nobody's talking about it. You know, it was a hype. They tried it, it didn't work." But if you really look at the talks of the people talking about it, it's all of the platform and SRE best practices just applied, uh, with u- with agents in it, you know. 'Cause if you have a solid platform engineering practice, you have proper, um, you know- Yeah. I, I, I mean- ... set up for CI/CD, you're gonna be faster with agents. Yeah. Listen, I'm, I'm interpreting what you're saying, like it's if, if you adopt one of these tacos, like, like if you, if, if you a- if your organization adopts Atlantis, you know, nevermind the AI stuff, you're, you're in a better position to do the work. Nah. You know how they always say like DevOps as a culture it's about the, it's not

[1:06:00] about the tools and so on. What I'm trying to say is if you adopt Atlantis, you're not gonna get it, you know, you're not gonna get an immediate- Okay ... improvement if you cannot adopt the culture and the abilities to, you know, rework your, the way of working to, to make, uh, you know, to, to get the benefit out of Atla-Atlantis. Like- Okay. So it's, it's- I mean, you could- Yeah. I am focusing on the tool, but... Yeah. I mean, you could, you could say that we both use a different definition of the word adopt because adopt, it's not just install the tool. It's adopt means we need to actually, um, put in the work of the, of the teams to be able to work with that tool and put in place the practices so that the tool can do the work, right? Um, so adopting the tool is, is can, can take a lot of effort. Uh, and, and only then- A little problem ... you can also have agents in the loop. I, I, I need to run in a bit, but I just wanted to explain one problem at work. I think you might find it funny. So we have consolidated a few projects into a monorepo around security

[1:07:00] engineering. But the problem is, is that now our, our security gate, which is quite, uh, pedantic, fails because like there's, there's a dependency in, you know, some bar project that's nothing to do with my project Failing my pipeline. So ba- so, so I had to make a pull request for Trivy to ignore someone else's CVE so that I could un-- So I could, uh, push out a change in, in my part of the monorepo. You know what I mean? So the monorepo- Mm-hmm ... the security tools don't seem to be aware that there's, like, different code owners in, in this project. And, you know, my, my, my role- No, the, the security tool doesn't seem to be aware that your dependency or declared dependency, um, is- Nothing ... like you're not really... Like, I mean, from what you just told me, I can understand it either way. Either one way is your project is using a base AMI that has a CVE,

[1:08:00] and that needs to be patched, but that you're not the owner of it. Yeah. In that case, it's good that Trivy is blocking you or not because you can't, you can't rebuild your, your, your project because the base AMI is not patched. Uh, so you can't redeploy even though it's running- Yeah, but, um, but to be honest, this, this, this... Okay, it's not an AMI, but this AMI has nothing to do with my delivery, by the way. So, so you're saying you're not even building on top of that AMI, it's just monorepo with like- Yeah, exactly. Exactly. It's, like, completely separate, but it's in the monorepo, and I'm blocked because Tri... And, and the way that Trivy does is, does is ignores, you just, like, put the CVE number. You don't say, like, "Ignore this path." Where. Ah. Oh, no, that sucks. Like, you're just saying, "We are accepting the risk of this CVE." You're not saying where. Hmm. Exactly. And, and it- SonarCloud and, and, and Aikido Security, you have to specify- Yeah ... in this file, ignore this, this CVE. Exactly. The, the... N- no security tool that I've seen, at least the, the one that I'm using this, in this client, uh, seem to be aware of the code owners. Yeah. So

[1:09:00] basically, it's all or nothing, and then it, and then everyone's getting blocked on each other because there's a few TypeScript projects inside, inside here with AWS CDK and all that sort of stuff, and there's, there's things ringing off every millisecond. But, but, but Kai, open weight models can hack your systems. You must... Security is a top priority. You can... Uh, it's like, you know, the red bell, the, the p- the, the- Yeah, but- ... the conveyor belt stopped. It's the emergency brake. It's about prioritizing- You have to go patch that system ... these gates. I mean, these gates are, are really painful because sometimes, sometimes you want to do a gradual incremental change, right? And the gates n- never seem to re- realize that what you have out there is, is already a problem. And you just wanna get the flow to fix things. Have you read this article? No. February 2020, uh, 2006, not even 2026.

[1:10:00] Continuous integration on a dollar a day. The reason I'm talking about the emergency brake, um, because in this thing, it's very simple. You have one computer that is used to deploy, and you get a rubber chicken, and then everybody can do what they want, but you automate your build. But if the build fails, if the build is red, you pull the emergency brake, everybody drops what they're doing, and you fix the CI/CD. And that's how you do continuous delivery. Um, I thought it was an interesting, uh, article. Uh, I actually started doing, uh- Well- ... DevOps in 2006, so ... it ma- it makes sense when you have a good ownership. But right now, since we've followed your good advice, Vincent, and gone monorepo, it's just a... It's insane now. Oh, so you were, you were complaining about monorepos, but what I hear is you adopted the monorepo without proper tooling. And it's not about the tools, but it is about the tools.

[1:11:00] Okay, let's end it here. There's a few things I need to catch up on. Otherwise, um- No, but the fact that your security system cannot bu- Like, um, it just sounds like your monorepo build is horrible. Like, that's why I, I use Turborepo. It's not about the tools, Vincent. It's about the culture. How many times do I have to tell you? No tools, only culture. Yeah, it- Oh, and, um, final note, we're going to Config Management Camp. I don't know if it's worth telling our audience be- because your talk got accepted and mine hasn't. No, mine is only an Ignite, I just realized. Well, it's still a talk, man. Jeez. No, they... Who, who goes to Ignites? I really- Everybody goes for, for lunch during the Ignite ... I like the lightning talks. I'm a lightning talk fan. I'm all for Ignite. Yeah. Um, but you got yours accepted. Mine bloody well didn't. Did you even submit an Ignite? Yeah, I did. Oh, okay. So- I guess, I guess my Ignite and my, my, my, like, 40-minute talk were very similar, and I guess they decided

[1:12:00] that's not worth 40 minutes. You get five. And I felt upset. Well, at least you- And I realized- ... at least you got, at least you got something, um, accepted ... I was more upset about the fact that they were right. And I was like- ... "Goddamn it, they're probably right. That's probably not worth 40 minutes." Okay. So you... Did you buy your tickets? We're going, right? I'm gonna buy my Eurostar tickets. Yeah, I have to... Uh, I, I have to go to FOSDEM as well. I wanna go to- Oh, really? You, you have some concerns about that. Um- I'm not too sure about going to FOSDEM, uh, because, just because I might not have time. Just because... You, we talked about it last time. But first, second, and third, we're gonna be in Ghent. So if anyone's listening- You mentioned this last time, and I said, "Are you sure you wanna know? What if somebody has a bone to pick?" You said Opus was good. I'll, I'll, I'll, like... No, no, no, it was Vincent that said it. You're misremembering. It was all about the- No, we're, we're, we're reading this script from Astra. It's all about the culture. Astra wrote this script, not us. Oh, yeah.

[1:13:00] Okay. See you. Bye. All right.