Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
March 30, 2026
·
Boston
CoChef: Your AI enabled copilot in the kitchen
Discover CoChef, an AI kitchen copilot using cameras and LLMs for real-time cooking advice, visual/auditory feedback, and interactive Q&A.
Overview
I am currently working on a device that provides users real time advice on their cooking. It is run on a Raspberry PI and uses a thermal + RBG camera to see what is happening in the kitchen. We have lights and speakers to give users visual and verbal feedback. There is an option for the users to interact with the device and ask any general questions.
Video
Transcript
Generated 3 months ago
Summary
Generating a talk summary...
View full transcript
Speaker 0: K. Thank you. K.
Speaker 1: Now Will do share.
Speaker 0: Okay.
Speaker 1: Hi, Sebastian. Alright. Pizza has arrived. If you guys wanna help yourself to the hot pizza, it's gonna be in the back. And there's also drinks in a in a fridge down the hall.
Speaker 0: How you doing? Good. Good to see you.
Speaker 1: Is this yours?
Speaker 0: No. I don't want.
Speaker 1: That's kinda like a trip hazard. Thanks, David. We can maybe open up couple of these just so they're easier. These are just the same as those. So we'll I guess we'll start with these, and then when these run off, we'll
Speaker 0: Why don't you just
Speaker 2: open out of the fridge?
Speaker 1: Yeah. Let's just leave them open.
Speaker 0: I think that way people grab
Speaker 1: it. Yeah. Otherwise, they won't know that in
Speaker 0: the fridge, it looks like.
Speaker 1: They don't wanna raid the fridge. Yeah. I'm just gonna grab 1 of those.
Speaker 0: Hey. Hey. Good. Yeah.
Speaker 2: Alright. Good. Good. How's your how's your business going?
Speaker 1: Well, now it feels like spring.
Speaker 2: That's weird. That's right.
Speaker 0: Well, once
Speaker 2: I saw you, it was like
Speaker 1: Yeah. So we had there were 2 a tinker's meetings that both got canceled.
Speaker 2: I couldn't make anyway, but I I didn't I had registered either of them, and then we'd run them out. So I
Speaker 0: was like, it's been a Will, but it's been a while. Yeah. Ian. This is Lee. Lee, 1 of
Speaker 2: the organizers. Meet you. So she's not here, like, invited to to work.
Speaker 0: Alright. You're welcome. First time? First time. Yeah.
Speaker 0: Hi, Mark. Hi. I'm on New Deal. Nate. I don't know how he got my guess.
Speaker 0: Or why? Alright. They didn't
Speaker 1: give us any plates, did they? This map oh, there are plates. There are plates. Here.
Speaker 0: More pizza. Maybe we should maybe we should put the the
Speaker 1: extra pizzas up here. We can put these up here just to, like, overflow. Thank you. I'm just gonna grab some.
Speaker 0: If it's saying you might not work in 1 third. Yeah? K. Hopefully, we don't run out of The check-in isn't more bangs that
Speaker 3: I care for. But it's showing that I wasn't
Speaker 0: Can you test it?
Speaker 2: Do web font.
Speaker 1: Okay.
Speaker 0: Yeah. I was trying to like, I got a client for the product, but I it shows that this the RSP is an unexpected. So it's a security. Yeah.
Speaker 1: Okay. Let me see if I can figure out what's going on with that.
Speaker 0: Yeah. I did not ghost this.
Speaker 1: You try that 1, see if that works.
Speaker 0: Thank you. Yes. So nice. Nice.
Speaker 1: Thanks for letting me know.
Speaker 4: Hi. It's Aliyah.
Speaker 0: Audio. Yes. You're Aliyah.
Speaker 1: You're Aliyah's friend?
Speaker 5: Alif. Mhmm.
Speaker 0: Nice to
Speaker 2: meet you.
Speaker 4: We brought you our ink today. So Oh, okay. I don't know if you probably, like, we were thinking to see if we can do, like, a community partnership for
Speaker 2: you guys. Okay.
Speaker 4: So they are here today just to see how everything's going and how the program is, like, that. Okay.
Speaker 0: So if
Speaker 4: you wanna meet them to them, we're gonna be into that. So I don't want it app’s, like, spend on this.
Speaker 1: Yeah. We're gonna we're just waiting for people to get their pizza and sit down. We'll probably start soon.
Speaker 0: Right. Well, I could run that. I remember the InterSystems. Thanks for coming. Hi, Nick.
Speaker 0: Hey, Wafa.
Speaker 1: Hi. Good.
Speaker 0: Hi, Will. Sorry. I got an x this morning, literally from 04:00 now I can nonstop, like, in this informal, like, getting into the car drive through. So cool. Okay.
Speaker 0: Well,
Speaker 1: you're here now. So, can you identify the speakers? I only know him. He's he's app’s speaker. So we're
Speaker 5: That's perfect.
Speaker 4: And And then the 3 of us?
Speaker 1: 3 of us, him, and then who's who's the other person that's And the
Speaker 0: I'm I have him.
Speaker 1: Oh, it's Jayesh.
Speaker 0: Oh, yeah.
Speaker 1: But I don't see him here.
Speaker 0: Yeah. I know.
Speaker 1: Oh, there he
Speaker 0: is. Yeah. Krish. Hey, Dave.
Speaker 1: How are you? You good to present tonight?
Speaker 6: Yes. I am.
Speaker 1: Okay. I don't know the exact order. We probably won't go actually oh, you're the first are you ready? Do you wanna go first?
Speaker 6: I'm more than happy to go first.
Speaker 1: Okay. Why don't you go first? We'll we'll go in this order except I think since my co presenters aren't here yet, we're gonna skip over me, and I'm gonna go at the end.
Speaker 2: Oh, who do you send anybody?
Speaker 1: Chris and Nikolai from Sunnica.
Speaker 6: Nice. Yeah.
Speaker 1: We're gonna we're gonna talk about Sunday clock.
Speaker 7: Clock if
Speaker 0: if they ever get here.
Speaker 1: That's pretty cool. Well, I'm gonna present no matter what that if they get here, they're gonna co present. Okay.
Speaker 0: They're probably just working on this. Hi. Hi. Yes. Nice to meet you.
Speaker 4: Nice to meet you as well. So who do you want to go first?
Speaker 0: Him? And
Speaker 1: then the other speaker is, like,
Speaker 0: the guy with a a black shirt.
Speaker 5: Yeah. Of
Speaker 1: course. And so we're gonna have a little microphone for you to
Speaker 0: use and, Will a Zoom meeting that you'll connect to.
Speaker 1: And then we also have, like, a little dongle for connecting to the and there's, like, some software you have to install the first time. It's kind of a pain.
Speaker 0: Oh, okay. But,
Speaker 6: connect to it now so I can, like,
Speaker 1: get it installed ahead of time? Yeah. It's called flick click something. So in order to use these projectors, they have, like, a dongle Will share. Yeah.
Speaker 1: If you can download the click share software
Speaker 0: You could just maybe just screen share. They could just okay. Yeah. Yeah. That's that's gonna be
Speaker 1: easiest for me to just share their screen.
Speaker 0: Yeah. Okay. Also, actually, I love the snow.
Speaker 2: The breeze really
Speaker 0: You love the what? The snow.
Speaker 1: Oh, yeah. I've I just duplicated the December, and it was there was snow in December. Hello? Yeah.
Speaker 0: I need.
Speaker 1: I need for you. I didn't hear you.
Speaker 0: So do you do you have a brief do that. These are the
Speaker 1: I'm gonna go last because I'm I'm waiting for 2 2 people to come. Hey. Hi, Sylvia. Nice to see you again. This is Wafaa.
Speaker 1: He's 1 of the co organizers. And this is Will. He's 1 of my well.
Speaker 0: Yeah. Will Who is that? Yeah.
Speaker 1: Yeah. Have have it grab grab some pizza and some drinks if you want and
Speaker 0: make it look How come I
Speaker 1: have not seen Chris.
Speaker 4: So I'll go third. You'll go fourth. You'll go fifth.
Speaker 1: I'm gonna go last.
Speaker 0: That's it. We have, I think, 5 speakers. 5 seats.
Speaker 1: Oh, you're doing it with Chris? Okay. Yeah. So you and I should go last, like, 4 or 5.
Speaker 5: And then I'm 3 and then
Speaker 0: 1245.
Speaker 1: 12345.
Speaker 0: Will, it's beyond. We should so you're you're saying that the Yeah. The new people organized. Yes. Of course.
Speaker 0: Where are you going? They're going first. Good. Yeah. Then I'm done now.
Speaker 0: I think I I think them app’s those 2 are done app’s them, it's pretty much just I get lost. They're just like, alright. Cool. Yeah. Because, you know, mine is technically too big.
Speaker 1: Yeah. Mine is kind of 2 presentations, which didn't happen too. So it's that's why I'm thinking it's gonna go last because
Speaker 4: I need a very simple. I like
Speaker 1: I had a hard time getting it into
Speaker 8: 1 thought. Because, you
Speaker 9: know I have
Speaker 0: I have proof, which is just yeah. This is just an enter button. Like, if you have a quick release and everything in the initial. Thanks. And then Justin has, Calvin,
Speaker 9: but she actually has, like, a small
Speaker 0: little thing to do. So Okay.
Speaker 9: It'll be, like, Will dose of,
Speaker 0: like, here's agent. Highlight this. Here's proof. And then, there's, you know, he'll Will Wafaa be a little
Speaker 4: That Will normally so you're gonna send us the same link?
Speaker 0: I would just
Speaker 1: Yeah. I'll have a little, piece of paper with the URL for the the code that they have to enter to join. Yep. And then we have these microphones. I'm wearing 1 right now.
Speaker 1: We have another microphone up app’s, and those are connected to the Zoom meeting so we can record it.
Speaker 0: Yep. EJ and I with the Yeah. The little gyro thing.
Speaker 1: Oh, yeah. I I have that. I I haven't really used it because that requires someone walking around the room holding it.
Speaker 0: Sorry?
Speaker 1: I haven't met anyone.
Speaker 0: You just ask them what's going on.
Speaker 1: Papa, I'm gonna invite you to the the meeting.
Speaker 4: Sure. Yeah. And do you want
Speaker 5: the invite actually the
Speaker 1: you're not in
Speaker 0: my contacts. But do you This Wi Fi.
Speaker 1: The Will Fi is 3 3 5 5.
Speaker 0: Yeah.
Speaker 1: This click share thing, it's This is what it looks like. Click click share. You can click.
Speaker 0: I like going to You joined the Zoom, and I'll I'll
Speaker 1: just we'll just have this computer always and that way people don't have to try to try to get this software installed and get this.
Speaker 0: I'm
Speaker 1: gonna go grab another piece of pizza.
Speaker 0: Good.
Speaker 1: Easy pepperoni. Easy pepperoni seems to be the exciting flavors.
Speaker 0: Yes. Okay. It should be good at both. You too? You have it on your Yes.
Speaker 0: I'm like
Speaker 1: I might have to admit them to the room because it doesn't I can make you a co host if you're
Speaker 0: Oh, yeah.
Speaker 4: Yeah. Make my account. Did you have to admit me? Yeah.
Speaker 0: So
Speaker 1: when people join, you'll have
Speaker 0: to admit them before they can.
Speaker 1: Hi, Superteam. Good. How are you?
Speaker 0: So sorry about that.
Speaker 1: Still Will bit of a tripwire.
Speaker 0: You got the notes for the video of the sessions. Right?
Speaker 1: Sorry?
Speaker 0: I see you got the notes of video of the sessions, or you've been doing that in a Will.
Speaker 1: We've been doing it for a while. Yeah. This 1 empty or is
Speaker 0: there stuff something so Oh, yes. Slow and paper. Will,
Speaker 1: we've actually ordered less food than we than we usually do. Because we are like
Speaker 0: I can get a Will too. Yeah.
Speaker 1: What time
Speaker 5: do you wanna start?
Speaker 1: 06:30. You wanna start now? Might as well we know might as well start now.
Speaker 0: Do you want me to time people?
Speaker 4: No. It's fine.
Speaker 3: I think it's good I think it's good
Speaker 0: to time her out.
Speaker 1: Do you think you need a chair up here? Is it people should people should be standing right when they're like, it'd be nice if this were a little taller because, like, it's I don't know if there's if there's, like, a stand that we can put the laptops on. Pizza boxes. Yeah. Grab some pizza boxes.
Speaker 1: Perfect.
Speaker 0: Hey, man. How you doing?
Speaker 10: Doing good.
Speaker 0: Good to see you. Sweet. You guys guys, feeling good? Yeah. Nikolai?
Speaker 1: Hey. Hey, This is our our very fancy laptop, Sam.
Speaker 0: K. You wanna place that and get pizza. Warm or is it cold? I'm sick.
Speaker 1: Go go grab some pizza. I'll pack.
Speaker 5: I think,
Speaker 0: the ages maybe can ultimate and then so, yeah, like
Speaker 2: the it's good.
Speaker 0: Do you think that's good?
Speaker 1: Well, they're they're gonna put their laptop here. This this 1 can be over here.
Speaker 7: It's the I
Speaker 2: think it's
Speaker 0: that 1 can get another part. Not that
Speaker 1: Depends on how tall they are.
Speaker 0: So I think it's 2 more. Yeah. Well, it Will be an adjustable laptop, Sam.
Speaker 1: Good. Yeah.
Speaker 0: Okay. You need me to handle it. Okay. Thanks.
Speaker 1: So last? We're do we're gonna go last. Yeah.
Speaker 0: I fixed the problem.
Speaker 3: You did? Well, I asked something called to fix fix it, but as a result, it takes much longer to
Speaker 0: do. Yeah. So maybe I will revert it back.
Speaker 1: I I also have a Telegram log of it doing it from before, so I could I could show people.
Speaker 0: Yeah. I wanted to ask people what they want to build Mhmm. Before you start talking.
Speaker 1: Oh. Then And you'll start it off.
Speaker 0: Yeah. Yeah. Start building. Cool. And when I'm, like, you will stop talking Yeah.
Speaker 0: Will. Alright. Cool. Alright. Thank you.
Speaker 0: Alright.
Speaker 1: Will get started. Okay. If I could have your attention, we're gonna get started. What? Well, this is this is only for recordings.
Speaker 1: This isn't a can everyone hear me okay in the back? Okay. Welcome, everyone. I heard there was some red line issues, so you all made it. Good to see everyone here.
Speaker 1: How many people, here is their first time? Wow. Like, half the room. Okay. Well, welcome to everyone.
Speaker 1: For those of you who for whom it's this your first time, we do have a check-in procedure. So this QR code tells us that you were here and we keep track. So, I encourage you to check-in that you are here so so we know that. Also, if you are hiring or looking for a job, we also have a a job board. You can scan that QR code down there.
Speaker 1: But a few logistics things, the bathroom is just down the hall. There's also drinks in the fridge and help yourself to pizza if you haven't already. I'd like to thank our sponsor for this venue, DXP. This is our first time using this app’s, so we're still kind of getting the lay of the land with the projector and the Will Fi and the the food and everything. So if it feels a little bit discombobulated, it's because we're still adjusting to this new space.
Speaker 1: But this is a great location. Right? It's like right in Kendall. It's easy to get to. They're not charging us anything for this space.
Speaker 1: It's also amazing. So, I don't know if are there any BXP people here? I think they may have already left. Well, anyways, let's give a round of applause to BXP.
Speaker 0: Thank you, BXP.
Speaker 1: So our agenda for tonight, we're gonna we've already done this. We're gonna start talks in just a few minutes. This is our lineup. We have 5 speakers tonight. There were a lot of people who wanted to talk tonight.
Speaker 1: We had 10 submissions, so that means we had to turn away half the speakers. Not because the talks were bad, but just because we had more people who wanted to speak than we had, slots available. So if you did submit a talk and it wasn't accepted, I encourage you to, submit it next month and, you know, hopefully, we can accept you then. We also had over 200 people register for the event tonight, and, obviously, there's not that many people here. We had to put a lot of those people on the wait list.
Speaker 1: So if you're here, it's because you've been vetted by our esteemed AI tinkerers committee. So, consider yourself lucky that that that you're here, and I encourage you to to share the link with your friends and bring people with you next time. I should mention our team. So Wafaa is 1 of our organizers, as is where's Will? As is Will and there's app’s other folks who aren't here tonight.
Speaker 1: Hi, Andy. Welcome. If you do wanna speak in a future meetup, we encourage you to go to the website. You can submit a talk and submit it. And there's a really cool, like, AI evaluator tool that will give you feedback on your talk and help you make it better.
Speaker 1: Like I said, we're reviewing those every month to identify really interesting talks. So we encourage people to show something that you're building. So we are all tinkerers here. That means we want you to show us about the struggles, the the takeaways, the learnings that you've, you've made from building stuff using AI. So we're not looking for PowerPoint presentations or look at my cool products like there's other app’s for that.
Speaker 1: That's not this meetup. This meetup is for people who are experimenting. They're learning the the latest and greatest AI technologies, and we wanna hear about it. So if that's you, feel free to submit a talk for the next the next meeting. We do meet every month.
Speaker 1: So, sadly, we had to cancel January and February because the meeting just happened to coincide with the snowstorm. So we're very happy to be back in session. And I think with that oh, 1 last thing. If you would like to sponsor us, tonight, the pizza is sponsored by Google from, like, a couple of app’s ago. If your company would like to sponsor, you get your logo up here, you get it on our website, you get to come up and say a little bit about your company and what you do.
Speaker 1: So, we encourage you if that opportunity sounds interesting, come talk to me or Wafaa or Will. Okay. And with that, I think we will get started. Our first speaker tonight is gonna be Jayesh. Come on up, Jayesh.
Speaker 1: Let's give him a round of applause. Yeah. You can put your laptop on this very fancy laptop stand. We're tinkerers, so we're very resourceful.
Speaker 6: Oh, Okay.
Speaker 1: I just sent yeah. Send a request. And if you don't mind, ring that.
Speaker 6: Check. Check. Can everyone hear me?
Speaker 1: Well, this is only for the recording. Oh. It's not actually I think if you just project, everyone should be able to hear you.
Speaker 6: Perfect. Alright. Will I think it should be
Speaker 0: sharing.
Speaker 1: Does it Sharing your whole screen? Or
Speaker 6: Yeah. Let me double check. I should be sharing my whole and I I did click desktop. You'd see a terminal? Sharing shares.
Speaker 6: I think you're all yeah. You're also
Speaker 1: I have to stop sharing.
Speaker 0: Yep.
Speaker 6: Yep. Perfect. So just a bit of background. I was supposed to demo my drone that was voice controlled, but during testing, it crashed a few times. I was I was telling it stuff like fly around in a circle to no.
Speaker 6: I didn't tell it to crash. It's just a byproduct of bad coding. So I'm demoing something else. So I've been obsessed with this idea of, like, LLMs and vision because I think 1 of the really cool with LLMs, like, we can do a lot of cool applications. Now some things that can really understand human intent and combine it with visual understanding of what's going around.
Speaker 6: So this is a device I've been building for the kitchen called, Co Chef. And I recorded a video last night, and I'll show you how the logic of the application works. So this is a device that's sitting on the range hood looking underneath, and, hopefully, it will prompt, what are you cooking today? Scrambled eggs. So what I'm doing here is that we wanna make this InterSystems, so it should understand what a user is trying to do in order to give it feedback.
Speaker 6: So it has a thermal so this is running on a Raspberry Pi. It's connected to a thermal camera, an RGB camera. It's so it's able to look down, and it's able to understand, you know, the temperature and other details. And so, unfortunately, the sound is not outputting. But if you see at the very bottom, there was, like, text in orange that was popping up.
Speaker 6: So this is actual audio that's been given back feedback back to the user.
Speaker 4: Scrambled eggs need low medium heat.
Speaker 0: Oops. Got it. Got it. Add butter now.
Speaker 6: Yeah. So it's able to see that, you know, like, it's telling the user, oh, you know, you gotta do this. You gotta do that. So what it's doing is that it saw that I wasn't scrambling my scrambled eggs. So it's saying, hey.
Speaker 6: You gotta do that. So you it's, you know, it's able to give positive reinforcement feedback too. And, you know, the thermal camera is also there to make sure that things aren't burning. And, unfortunately, I couldn't record it in this video, but there's also lights on the device. So not only is the user getting, like, audio feedback, but it has lights that have different colors of to flag the different stages.
Speaker 6: And, so we also have different modes like cooking mode, done mode. And now it's telling the user, hey. It's too hot, so you better, turn off the burner. And, yeah, that's it. And I'll show you a bit more about how the device looks like.
Speaker 6: So this is what it looks like. This is how it looks like on a stove, and this is how the lights look. Like, so this we have alerts. So a lot I've had a lot of issues. This demo looks really nice.
Speaker 6: I'm glad to show it. It does not work for anything more complicated at the moment. I tried cooking chicken, and it kept yelling at me add oil. So I was like, okay. I'll add oil.
Speaker 6: And then it couldn't see that I added Will, so it kept yelling every 5 minutes, add oil, add oil, add oil. And I was like, oh god. And then my chicken was white, but for some reason, it thought it was overcooked. So what happened is that I realized I'm hitting and I'm using Opus 4.6, which is a really good model. I thought Gemini 3 would do better because that's what people use for, like, visual based analysis.
Speaker 6: It did not do that great, and it was terrible at tool calling. So some of the improvements I need to make is, okay, you have an agent. Now what it needs to be able to do, it needs to be able to call computer vision models, like Panoptic segmentation, for instance, in order to be able to understand, like, okay. If you have a stove and you have multiple pots, which pot do I need to focus on? What are the individual ingredients, elements of the thing that has been added to the dish?
Speaker 6: Like, for example, 1 piece of chicken might be undercooked. So the panoptic segmentation can highlight that piece of chicken. The LM can use as a tool and then be like, hey. This 1 piece of chicken is undercooked. You need to turn it around.
Speaker 6: There's also the heat data. The way that I'm passing the heat data is it's like a NumPy. So instead of passing in the entire set of numbers, I'm just passing it the delta. So I'm passing this is the hottest, this is the coolest, this is the bits where there's the biggest temperature difference. I have no clue if you're gonna be able to tell good temperature data if you're trying to boil an egg, for instance.
Speaker 6: So a lot of these are things that I'm working on and need to consider, but this was, like, a proof of concept that I built over a weekend. And, what's neat is that this is running on a Raspberry Pi. So, like, right now, I'm doing just a I'm using a open router to, like, call an LLM. But I'm thinking down the line, maybe there's, like, some models I could put on the Pi so it can work a lot more efficiently. Yeah.
Speaker 6: This is my demo. Yes. Yes. Yes. For now, I'm just using a screenshot and sending it to the LM, but the goal would be we don't wanna use the LM because it's also very expensive.
Speaker 6: So we wanna have some sort of, like, smart thing where you use a vision model. And the main thing that I think really helps build products like this is this concept of states, understanding what is these the state transitions that you wanna consider. For example, there's no point calling it LLM if the vision model detects it. So the same things have been there for a while. Only when there's a change, you see some movement.
Speaker 6: That's when you wanna call the LM. But yeah.
Speaker 1: Yes.
Speaker 6: It's just that thermal sensor and, an RGB. So I guess what it meant was, like, for example, this is, like, if you're trying to do a boiled egg, I'm not sure. For example, if multiple boiled eggs in a hot pan and a user specifies that 1, like, easier like that. Like, I'm not sure, for instance, like, do we need to be able to look at the temperature inside the egg or versus something like cooking meat. Right?
Speaker 6: Like, Will cooking meat, you really need to know some of the inside temperature. This sort of device, looking at it, I very difficult from the outside to see if the meat's done or not. So these are some of the things we have to consider. Like, how would we solve this if this is something that can't be solved?
Speaker 0: You say say, but, Yeah.
Speaker 6: We probably can, but I think we wanted, like we didn't want like, ideally, if you're gonna sell this, you don't Wafaa use or have to also have to buy a lot of these things if they don't care about it. But you're right. Like, with IoT, you wanna if they have other sensors, you wanna bring it all. Do you have a No.
Speaker 0: It yeah.
Speaker 6: So that's where we're, like, feeding it deltas. So we're telling it more about the some spots we're telling you the actual temperature, and then we're also giving it the change of temperature over time. So we want it to be wary, like, a few seconds ago, this this is it. It was this. This is what it is now.
Speaker 6: But, also, like, the LM isn't smart sometimes, so it's like it thinks, like, oh, it's ready to cook at about 50 centigrade, and that's not very helpful. Do you wanna start cooking at 70? So there's, like we have to be careful with the prompting, prime different models, potentially building a rag to, like, solve this. Yes. Sorry.
Speaker 6: You there with the I don't know names. No. Because inference is really inference by far is the most important thing. So we also even had to, like Helm's app’s so freaking verbose. It took, like, 2 or 3 prompts to get it to be like, if you're giving users feedback, be as snappy as possible.
Speaker 6: And 1 of the things that we really have to implement is this idea of prediction. If you're gonna give users feedback, you need to give it to them 5 seconds before the thing is gonna be done. Because you can't tell a user this is done the moment it's done because it's also, like, the inference, those people are slow to react, and so forth. Yes.
Speaker 2: So, this is really cool, by the way. Really, really really awesome.
Speaker 0: Thank you. So kind of
Speaker 2: 2 parts. 1, I just wanna confirm. So the idea is like a a cooking coach.
Speaker 6: Yes.
Speaker 2: Kinda idea. Right? Like, you threw this device to help me make something.
Speaker 6: Yes.
Speaker 2: So I I think that's that's really InterSystems. But, but it's kinda similar to infrared. I've even I like really limited data that the LMC today. But there's, like, tons of, like, cooking shows with narrative. Yeah.
Speaker 0: Have you
Speaker 2: thought about custom models or fine tuning?
Speaker 6: We probably would have to do that. We're not at a stage yet where we're ready to fine tune because we are the stage, right, where we need to build up the infrastructure to best of our ability before we're gonna see performance gains from fine tuning because it's also, like, a very long process. And, you know, I'm glad you brought up cooking shows. 1 of the reasons I've actually built this is because I try to follow TikTok recipes, and they're the most unhelpful thing ever. I have to pause.
Speaker 6: It's a 1 minute video. I have to pause it, like, every few minutes, and then it says, yeah. Beat this until, like, custardy consistency. I'm like, what does that mean? So, like, 1 of the reasons why I wanted this is, like, if I ask, hey.
Speaker 6: Is this custardy consistency? I'm hoping an alum would be smart enough to know this is is it or is it not? But only Will probably achieve this with fine tuning.
Speaker 0: Yes. So
Speaker 6: it's a bit of both. I think for, I think, like, if it's an empty stovetop, that's like before you start the over the loop of going into the LOM, it would just be a vision model, and you would just be detecting, has there been any change in the environment? Like, has there any been pans detected? Like, even, like, the thermometer temperature. Right?
Speaker 6: If the temperature is not on and the user isn't asking any questions, probably no one's cooking active, things like that. With the LM itself, if the LM wants to know what's going on or, like, dig deeper into the individual pieces, that's when I would use, like, a LM being able to call a vision model to get feedback. Even, like, 1 of the things that's interesting here is that if you notice that this stove, there's you see 5 burners. In the demo I gave in my home kitchen, I have this many stoves. But because of the angles, like, it during the video demo, it only focused on 1.
Speaker 6: So So 1 of the things that we need to be able to do is get cameras that can zoom in and out. So you need a l m that's smart enough to understand, like, am I getting the full picture of the stove or not?
Speaker 5: But let me take 1 last question, please.
Speaker 6: Okay. Yes.
Speaker 1: Yep.
Speaker 6: Yep. So for the first part, this is, we had 3 d printed this case. My partner had. It's not being taught enough where it starts melting or start worrying to be melting, but you're right. This could be a concern.
Speaker 6: No good answer other than using a better 3 d printed material or something else for the case. Secondly, right now, we are doing every 5 seconds, which is, like, we just needed to get something out of the door. But what we would do in the future is this concept of state diagrams because we're like, the temperature sensor, Like, we would probably want, like, maybe a visual model temperature sensor to detect how demonstrates have transitioned. Like, has it become too hot? Does it look like there's a bit of browning?
Speaker 6: Something like that. Unclear, but this is, like, the logic we would use.
Speaker 5: Yeah. Thank you so much.
Speaker 1: Yes. And if you didn't get your question answered, we are gonna have a networking time at the end, so you can track down and ask them whatever questions you have.
Speaker 10: So hi, guys. My name is Sohan. I'm gonna show something called ATX, which is a short form for agent access. As you can see on the screen right now, it is a CLI tool Will works together with Cloud Code right now. And
Speaker 0: oh, yeah.
Speaker 10: Oh, yeah. I can make the phone bigger. Right? Yeah. So no.
Speaker 10: This is just 1 part of it. You don't have to stare at it for long. So this is just, it's a App’s tool that is working with Cloud Code right now and piggybacking its hardness stuff on Cloud Code. So, it's using Cloud Code to run it, and it is running with Cloud Code. What it's doing is, first part of it is you initialize sort of a you you pick any repository you want you're working in.
Speaker 10: You initialize ATX in that repository. What it does, it goes through your repository, looks at the different files you have. If you have CloudMD, if you have your AgentMD, it looks at those file. It looks at your commit history or committing patterns, how you are, committing information or how you're committing, code. And according to that, it writes up, review agents or specialist agents here.
Speaker 10: And at the same time, adds a MCP and bunch of folks to your cloud code. So what it does is it's kind of a continuous integration for your Cloud Code instance. So let's just, I'll give you an example. Right now I don't have a Cloud Code session running. If I try to do that, it's going to take a lot, a lot of time to actually make a change.
Speaker 10: And then for the loop of ATX to run. So I have, I I was doing testing before, and, let's just pick any 1 of the commits I did. What happens is, let's just say your cloud code will start working as it does normally. Will do its changes. It will go through this whole loop.
Speaker 10: And once it's done looking for its whole loop, there's a hook that gets added when you initialize ATX in your repository. And that hooks, sends what files has been changed in that specific session by that specific agent, cloud agent. Then there's a leader. So that's a leader and, worker type of thing where leader is orchestrating stuff between all of its workers. And, it's sending out files or so leader decides first which, which part of the change needs to be looked by which agent.
Speaker 10: So it decides that. It starts doing routing. And then, once it done done the routing, it goes to different agents. Different agents try to, review it. So sort of a continuous integration like how you run tests, but during your cloud code session.
Speaker 10: And then you get a response. And then according to that response, the, cloud session continues fixing it. It's not a loop. It will only do it once. We are not making it loop, but there is an option to make it loop.
Speaker 10: So it becomes a Raalf loop in a way, but then, there's also it's also LLM as a judge kind of, or a bug bot. And the other thing where which is there in ATX is also like bug bot. You can I tried it on OpenCodes repository? I gave it a PR, it, so internally whenever a PR links come, PR link comes into a plot code or through the ATXMCP, it gets converted automatically into a diff. It makes a work tree separately and checks it out so that the agents can go look at the whole file and, get the diff or make a review more holistically.
Speaker 10: So it shows up and it ends up giving a review accordingly. And I think just to show what kind of agent it makes, this is the open code repository. That's the, so when I entered agents here, it, went ahead and made, app UI reviewer, effective service reviewer, then product out, and session agent reviewer. And obviously, there's a test coverage agent. So, you can go ahead and see what the system prompt is for that specific agent.
Speaker 10: And at the same time, you can select, which anthropic model the agent uses. It's completely on Cloud Code right now because, it's it was Cloud Code CLI was the easiest to hack together, because of my, I'll I've written this in Golang. And because of that preference, there were no SDKs by any of the providers. So I had to hack it together and, wrap make a wrapper over the CLI itself. So, this is where it is at right now.
Speaker 10: So the future of this would be that, you can have any kind of model come into get plugged to ATX, and you can have your local Sergeant. And those act as a, CI reviewer or a continuous integration for your coding harness. So the dashboard, if if when when you have used it a lot, the dashboard shows you all the costs, the tokens you have used, and the output token and which model you have used the most. But, yeah, that's it. That's ATX.
Speaker 10: Okay.
Speaker 3: When you're doing your tool calls, you said that you're only on CloudQuote. Yeah.
Speaker 0: Just in a TypeScript SDK, but you were just in Go. Yeah. And you're
Speaker 3: just running a single, like, CLI.
Speaker 10: It's not a single session. It's multiple sessions. And, whenever so when you start a tool Will, let's just switch over to 1 of the reviews to show.
Speaker 0: So I noticed you had a tool call Yes.
Speaker 1: On your timeline. I'm curious how you
Speaker 10: It is. So it this is this is a single session. And when it's calling other agents, all of these agents are responding to its session of its own. And, there is a bit of back end work Will is done for those agents to talk to each other. If you see, there is 1 review.
Speaker 10: I remember. I'm just checking. Sometimes these agents try to answer the questions by themselves. There's a lot of prompting that has gone in here, the to make sure that the agent itself does not call it. I think this 1 did.
Speaker 10: And then I yeah. So as you can see, leader itself tried to answer the question here, but, I was then there is a lot of prompting that goes in on the back that is telling, okay, you can't do this. You have to call all the, all your specializations and only then reply. But it's happening all in 1 session. Because it's wrapped over the CLI, I have
Speaker 2: to do that and not the type of
Speaker 10: speaker. Yeah.
Speaker 0: Can you say a little bit
Speaker 2: about the the problem you're trying to solve and, like, maybe it's like why you kinda went down this road as opposed to using Will in agents and Claude and steal the whatnot?
Speaker 10: Yeah. So, 1 big problem with so I started this before the built in agent started with cloud. And but 1 thing I've seen with cloud built in agents is that they hog tokens a lot. They use a lot and a lot of tokens. I have tried using them when in Swarm initially came in.
Speaker 10: I tried to use it for review or even the new plot review thing. They ended up finishing my session Will, really quickly. You've seen that because of certain prompting styles and certain prompts that are there in the system, it's able to use lesser and lesser tokens, but give a more holistic answer if not, if not better. So that's why this makes it better in token usage. So the cost is like not as much as if you had given the same thing to Claude and try try to make it a agent's phone.
Speaker 10: That's why this route and plus, I think Claude does the same thing, but it also does the leader and specialist agent things. This this specifically, I'm not sure how they write the prompt for app’s, specialist agents. But as I said at the start, this uses your commits, your agent MD, cloud MDs, and takes everything into consideration to make those agents. So your coding patterns are Will to show up in those agents and the prod might or might not do it. It might give a generic more generic answer to your question.
Speaker 10: Yeah. Oh, there's 1 more. Yeah. This. Yeah.
Speaker 10: Yeah. For for some of it's the context size. When I started, the 1,000,000 models are not there. And I've seen that as you keep using the 1,000,000 context model, as you go past 200 k in those models also, decay starts happening and models start forgetting a lot of things. Right now, what you're seeing, as I said, each session is getting, each agent is getting spawned into its own session.
Speaker 10: So it's kind of getting its own context in a way. So that, in that way, there is no decay, as per so the agents are better at following those rules. Before we had this, when initially Cloud Code came in and when we are trying to do reviews with Cloud itself before when we are initially testing bug bot also. So that's the problem that kept happening. If you use the same session, what used to end up happening is Cloud used to give very random answers or very hallucinated answers because of that decay.
Speaker 10: So this approach saves you from doing that, and at the same time, the, token and cost usage is reduced from what I've tested. Mhmm.
Speaker 0: Yeah.
Speaker 10: Yeah. So, the the plot has the, has a hierarchy in in the folders. I forgot the hierarchy, but it does dot cloud slash session. In the same way, we do, dot a t x slash session, and everything is collected there, just to make sure that nothing is leaked across or nothing is missed missed across those sessions.
Speaker 5: 1 last question? Yes.
Speaker 0: Mhmm.
Speaker 2: Yeah.
Speaker 10: This itself by by some provider? How's 100.
Speaker 0: But it sounds like a lot of
Speaker 10: Yeah. No. So I think, now more and more people are adapting to it. As you said, everyone's doing their rendition of this, specific thing of agents and how you, make agents work properly. Right now, my conviction is obviously this this the app’s system of leader and orchestrator is working out very well.
Speaker 10: There are definitely more researches coming out beyond that. Maybe I might have my own research coming through this and findings that might get adopted by someone. So it might take time, but, at in today's age, you'd never know. It might happen next month. It might take 5 months.
Speaker 10: Yeah.
Speaker 5: Thank you so much.
Speaker 10: Okay, guys.
Speaker 0: Okay. Oh, sorry.
Speaker 3: Get rid of this. Yeah. I don't know.
Speaker 5: I guess you have to go see me. Alright. So, my name is Wafaa, and what I'm going to show you, it's, me graduating from lovable to cloud code. So it's a very proud moment for me to fully productions. So, basically, I'm try I'm training for, triathlon in June.
Speaker 5: And I was looking into all the different apps, and I was like, I could build this myself. And so, basically, I I will show you my what I built and what I'm using right now. But, so this is, like, user comes in and just they share their information. So once you sign in oh, password didn't match. 1 second.
Speaker 5: Some errors.
Speaker 0: Oh. Yep. Alright.
Speaker 5: So let me just sign into my account. So, basically, as Will part of the workflow so when the users try to sign up, it asks them, like, when when is the when is the race date? And then also try to get some information related to, like, what's your baseline. So, like, how much do you run, how much do you swim, how do you how much do you bike, and what's your benchmark. So based on this information, it's trying to customize your experience.
Speaker 5: So, basically, this is like so I have 68 days to go, and I have 10 weeks to go. And, basically, each week, it breaks down into, and in the onboarding process, it also asks you, like, how many days do you train and how much time do you have to train. And, basically, it customize your training plan so you don't have to think about it. So, basically, every week, I I have, like, my weekly training plan. And if I click on it, it basically break down, my training, what I need to do as a warm up, as a as a mindset, and also as a cool down, and also gives me some coaching tips.
Speaker 5: And, also, I could, like, switch between workout if I'm not feeling like I want to go for a run, but I wanna go for swim. I can switch and plan my my week accordingly. And what's really cool about this feature is as I'm checking in and saying, like, how based on my how I'm feeling, so when I do check-in and I say, like, I went for a run and how what's my if I'm tired or I feel really great, like, how much energy it took off me, how much my my sleep, it basically adjusts my plan, based on where I am in my my training journey. So each week, I get, like, plan change. And also I have my AI coach.
Speaker 5: So the AI coach connect with the plan and it's connect also with my check ins. So if I ask it any question, it's Will be like, oh, last week you run this 3 miles and you were dying. Maybe this week, try to just run 2 miles and take like, do this base because based on your historical data. And here I have my dashboard. So the dashboard is just really I'm a big dashboard fan, so it just shows me what I've been doing and how my, how's this week performance comparing to that to last week and how I'm how much I'm ready for my training.
Speaker 5: And then also I can put my benchmark. So I haven't put my benchmark yet because I'm still figuring that out. Like, still really training for it. But, yeah, basically, what was really cool about building this product so I Will it first with Lovable as a front end just to understand, like, what's my workflow is gonna look like. And I also use cloud skills.
Speaker 5: So I so I'm not a technical. So I have I'm a product manager. So I use cloud skills, the CTO to be to think with me through the problem and what what I need to build. And then I had my first prototype in Lovable, and I made in Lovable. I asked Lovable to create for me, like, a cloud dot m d and spec dot m d.
Speaker 5: And I took all this, put it in GitHub, and then I deployed everything into a cloud code, and I start building for the whole weekend. So 1 thing I learned through this process, like, with cloud code, I I use superpowers agents, and I was running out of credits. Like, by the minute I I say cloud, it tells me I reached my limits. So I paused that, and I start using more, like, the age like, the banner agents just to review to review where we are and what we need to do next. And something was very helpful for me is to clear my chats every every time I finish 1 features.
Speaker 5: I, like, test it on my local. If I'm happy with it, I deploy it into productions, and I go back into what I need to build next. So, yeah, right now I've been using this product, I think, for 4 weeks, and I probably maybe have 5 users. It's just my friends and I were training together, so we Will that. Yep.
Speaker 5: That's it.
Speaker 4: Yep.
Speaker 5: Yes. Thank you. I used lovable. Yeah. To do the to do the back end.
Speaker 5: Yeah. Yep. Yeah. That's, like, a feature too. It's coming app’s because I have Garmin, and I wanna connect right away, like, my data through Garmin instead of me doing check-in.
Speaker 5: I think this is more complicated, so I'm still figuring out the API connection and stuff, but it's I wanna, like, figure this out. Yep. Yeah. So, basically, here, just my check-in. So there's no calories here.
Speaker 5: So it's just basically, like, the the day, the type of exercise I did, and then also, like, how what's my durations, like, how long I went, What's my distance? What what's my effort? Like, how much I would effort into the exercise? Usually, I would 4 or less unless I'm with someone else. Then I I get my ass kicked.
Speaker 5: And then my sleep and energy. So there's no calories here. And then yeah. So that's, it's it's only on my watch. Yep.
Speaker 0: Because I know that if you just provide code database or a dashboard or app like that, I want them to work very nicely. The API works so so and the database is gonna pay off.
Speaker 5: Yeah. So, basically, this is why I use my planning session with cloud skills, like, with the CTO and, like, make sure that, like, I know what what data I'm capturing and where this data will live. And then Will was in cloud code before building the database. I was brainstorming first with with, with Cloud Code to be like, what type of data we need, how this data is gonna be moving because it's gonna be like, the user information is connected to the plan, and it's connected to the to the coach. And wherever you check-in, everything is connected.
Speaker 5: So I just, like, brainstorm at first with, Claude's Will, and then was in the Claude code. I brainstorm and, like, then Claude starts, like, saying, like, okay. Like, give me your, login and you whatever. Like, they ask for a few information from a SuperBase, and I use Super SuperBase as my back end. Yeah.
Speaker 5: It's it took few trial and error, but yeah. Yep.
Speaker 0: Like, expertise or knowledge that that you talk more about that?
Speaker 5: Yeah. So, basically, the coach so this is more I use, like, system prompt. So I prompt the the API to be my, to be a coach. So I said you're you're an expert in x y z. This is what I want you to do.
Speaker 5: This is what you need to say, and this is, like, the guardrail. Like, they should not be, for example, making things up if I did not exercise. Like, so I just put some guardrail, and I I tell it what type of personality the AI coach should have, which is, like, related to the sports I'm training for. Yep. That sorry.
Speaker 5: Meal. Yes. So that was my first app able to it's lovable only, like, to, like so, basically, I train and I I log into my what I'm eating, and it basically helps me with so I had that in a previous version, but not in this 1. Yeah. Well, thank you so much.
Speaker 0: Will?
Speaker 1: Will all of you nontechnical folks in the room We
Speaker 5: can't do it.
Speaker 1: Wafaa is not an engineer, and she built that. So very impressive.
Speaker 0: I can do it. You can do it.
Speaker 1: Is Chris gonna be upcoming?
Speaker 0: Yeah. We'll play it by ear.
Speaker 1: Alright. Let's just get the screen share rolling.
Speaker 8: Share screen
Speaker 1: should be oh, let's just share the whole thing. Just YOLO. How's that doing? Oh, okay. Let's get out of this little workspace there.
Speaker 1: Alright. Cool. Alright. I think we're live. Oh.
Speaker 1: 0, that is yeah. You did that. Yep.
Speaker 9: Alright. Hello, everyone. My name is Will Sargent, and I'm the cofounder of Agentic Highway, which you'll see in a moment. How many of you have used Claude Code just by raise of hands? Cool.
Speaker 9: How many of you have used dangerously skipped permissions? All of your hands should be up. You've all used
Speaker 0: it. So,
Speaker 9: but do you really know what it's doing when it's running on your computer? Do you know what artifacts it's creating? Do you know what it's reading, and do you know what it's writing? And can you prove it? Well, that's what we're looking to do with Agentic Highway in our in our first kind of product here is, this is not a product demo.
Speaker 9: This is something that I've created that we're going to hopefully be, pushing out for free and and Will sorts of stuff. But, essentially, the way this is gonna work is that we'll do we'll do a we'll do a default scan here, and it what it will literally do is it'll scan your home directory on a on a Mac. This also works on Windows and Linux as well. And this is all written in Rust, and it'll go through, and it'll take a couple seconds to to scan everything. And we scroll up here.
Speaker 9: You can see I have a lot of stuff on this computer. But you essentially have all of those agents dot m d files that are in all those GitHub repos that you've cloned, all of those, all those MCP servers, all those agents, everything else that it finds on your computer, it pulls all those in and scans them for, are they executing code? Are they asking for network access? Are they doing dangerously skipped permissions? Are they doing shell shell execution?
Speaker 9: All those sorts of things. And so we're constantly building new detections in. Another neat thing about this, though, is that command line's not really that great for viewing information as you can see. So we have another thing, online platform called vetted. So if I hit yes on that, we'll go ahead and submit the scan to vetted.
Speaker 9: So and if we look here, this is vetted. So this is very much so so this piece is not public yet. This has a lot of bugs that I need to work out. But, essentially, this will allow you to take a look at all of those scans that you've, uploaded to the server. You can also view the scans on, App’s as well kinda locally.
Speaker 9: It just won't work quite as well, as, you know, a nice dashboard here. But, essentially, you'll be able to do prompts, skills, MCP servers, agents that and you'll be able to have an inventory of what you've got. You'll also be able to pull in from GitHub, from Bitbucket, and things like that. So this is very much still in kind of alpha beta at the moment, but, we're, you know, we're looking to kind of push forward on this soon. But yeah.
Speaker 9: So that's that's basically, what we've got going on with, proven vetted. We have another component, which I'll get to in a moment. But first, I wanted to pause, take questions, and that sort of thing.
Speaker 4: So I'm just really curious. After you do, the vetting and and figure out, like, the risk, do you, like, also, like, help the users take some actions against Will or this other product?
Speaker 9: Yeah. So coming soon. That's 1 of the things that we're that we're I was trying to do here was to say, how in a command line environment do I give the user as much information as possible without being overwhelming? The command line is not a great spot to give users information. And so, what you see here is essentially me telling Collab code, hey.
Speaker 9: Give the user, you know, the top 3 things that you found and then and then why they were concerning. And then really trying to push to, you know, upload those scans to vetted and say, okay. How can we then create an inventory for the user? How can we give them more insights into their data? Short answer to that is hashing.
Speaker 9: The long answer to that is more is in progress to make sure, because that's also a storage liability as well is to say what data what is the minimum amount of data that we need to store to give the user this result? Don't wanna store your entire contents, your hard drive up in the cloud. That's unnecessary and very risky. So we wanna store as little as possible while also being able to give you exactly the insights that you need there. I'll take 1 more question for now.
Speaker 9: Go over here.
Speaker 6: So,
Speaker 9: from OWASP, NIST, and a handful of other, you know, you've got, various different,
Speaker 0: I
Speaker 9: don't wanna say governing bodies, but, very various different regulations and whatnot that have come out to basically say, hey. Here's the OWASP top 10 for AI. Here's the OWASP top 10 for web applications. Here's the NIST, AI foundations and that sort of thing. So I'm basically essentially bringing those in and saying, okay.
Speaker 9: What can we easily look at from a deterministic standpoint? None of this is is nondeterministic. All of this is right now, it's all keyword matching. It's very much so we know something to be bad, so we're going to call it out. Because, you know, as much as I'd love to put all of this through a large language model, I really wanted to be as deterministic as possible with this first push out.
Speaker 9: So, speaking of nondeterminism, though, there was 1 thing going through all of app’s stuff that I realized, which was I didn't use a claw for any of this because I don't trust claws. And I realized I didn't trust any of them because I didn't build 1. And once I saw how kind of easy it was, not so easy, to build 1, I basically said, alright. Well, you should do something about that. And so, Justin over here, he's part of the agent of Kiowa squad, and he was like, hey.
Speaker 9: That sounds like a really cool idea. And I sat down, and I did the architecture behind it, and he will show oh, I oh, okay. I get him on Zoom. Crap. But, basically, the concept while they're getting that set up, the concept of Kelvin Claw is I don't wanna use a claw that I don't trust.
Speaker 9: Right? And you've got this whole clawcharts.com that shows you all of these different claws ranging from iron claw to open claw and everything in between. I don't have time to go looking through those code bases. So what I decided was, hey. Can I create the most minimal claw possible while also being extremely readable, extensible, safe, and secure so that you can just download it, have Opus say, hey?
Speaker 9: Here's what's going on with this claw. And and you can you can even read it for yourself and just go through the thousands of lines of code. This is I believe that last count, there was less than 15,000 lines across the entire core of it. So it's really designed to be simple, secure by default, and very, very minimal so that, you know, we don't have Will you know, we don't have Discord built in automatically. That's something you you would need to add in.
Speaker 9: Add in, like, maybe in a regulated environment. Sergeant don't need that. You just wanna be able to clock. Run the clock.
Speaker 7: That's,
Speaker 11: that's,
Speaker 1: how we
Speaker 0: do it. Almost.
Speaker 1: Alright. Almost done. Back to your screen.
Speaker 0: So Justin So
Speaker 6: Justin will go through
Speaker 9: how, how, how what the work he's done on Calvin Klein.
Speaker 0: Fair.
Speaker 6: And to be fair
Speaker 0: While this
Speaker 9: while this while while the architecture architecture and stuff was found out of
Speaker 3: out of mountain.
Speaker 9: Saying, hey. You know, this is a big problem. You clause is that you can't trust them. And they're just And they're just gonna go do crazy things. Justin.
Speaker 9: Justin basically app’s at and was like, hold my beer.
Speaker 7: Yep. Alright. Yep. Introduce myself again. I'm Justin Paciella, and I'm with the Genta Highway Group.
Speaker 7: And what I'm showing And what I'm showing you today is Kelvin Claw. Kelvin Claw. So We're gonna start off We're gonna start off with fun Interactive. Interactive live demo. So if you have Telegram so if you have Telegram give that QR give that QR code to scan and we're gonna play capture the flag your goal is your goal is to recover
Speaker 0: a hidden file
Speaker 11: a hidden file
Speaker 7: called flag dot called flag dot txt. Is he getting some,
Speaker 0: Is he
Speaker 7: getting some, traction going there? Traction going there.
Speaker 5: Will you try again?
Speaker 7: Flag dot TX. Dot TXT is what you're looking for.
Speaker 11: What else is going on?
Speaker 7: Alright. So what's happening here is I basically set this up in pairing mode.
Speaker 1: Like that?
Speaker 7: Yeah. Yeah. It's in yeah. It's in it's in it's basically this is set up with a normal pairing mode that allows me to pair this to my Telegram what I did is disabled that so anyone can message my claw agent so you guys are met so you guys are met talking to my personal agent and trying to recover Will that you're never gonna get you're never gonna get and here's why
Speaker 11: and here's why you
Speaker 7: can see from my
Speaker 0: You can see from
Speaker 7: my Telegram, I'll be able to tell it. Might take a sec because it might take a sec because there's a Pretty long queue. Pretty long queue of, tool calls here. Oh, we got some web we got some web searches. People are looking at the Wiki.
Speaker 7: Alright. So while we wait for So while we wait for this, to answer and finish app’s what's in the queue Anyone wanna ask? Does anyone wanna ask meantime questions?
Speaker 11: So what did you do differently with this claw? Did you did you, branch it from 1 of the existing ones of, or, like, how
Speaker 2: did you how did you get here?
Speaker 7: Nope. This is made from scratch. Largely by Largely by Will, and I basically want to step through most of it and debugged it. So What makes this different? What makes this different is deterministic measures.
Speaker 7: The the whole the whole theme of the of what we're building here is that the the measures The measures are don't rely on don't rely on how good your LLM is. So If we really wanted If we really wanted, you could hook this up to your I don't know. 0.8 I don't know. Point 8,000,000,000 OLAMA. OLAMA model.
Speaker 7: And Still have the same security. Will have the same security other than you wouldn't be sending Your data to your data to an LLM provider. Up. And now here's Up. And now here's the flag.
Speaker 7: Might be a little bit small, but What it says there what it says there. Prompt injection Prompt injection will not save you. Yeah. So this, so this, the demo. Demo of the for the for Kelvin clock.
Speaker 0: Have any Does
Speaker 7: anyone have any questions they wanna ask?
Speaker 11: So I have a question where so we might have heard of this 2 called ZeroCloud, PokerCloud, stuff like that. Right? Mhmm. So then you just put your key inside where you are specifically aiming app’s specific channel, and you are telling the thing that they just don't get anything apart from the stand. So why that approach is just out of the data, and what's your approach?
Speaker 12: Doing something different?
Speaker 7: Well, well, this approach this approach is fundamentally not our not our I think I might be misunderstanding you, but this approach is fundamentally no different from that. I basically just turn that off. Like, normally, you would not be using this with this turned off. Normally because, normally, I wouldn't want a bunch of people messaging my claw and And using it without using it without my permission. But I disabled But I disabled that disable that for the sake of the demo so that people could how it works.
Speaker 7: See how it works.
Speaker 11: Latency, like, I sent a message to about a minute. So because of all the articles that you had placed?
Speaker 7: That latency That latency would be because there's a lot of people in the room that are messaging it.
Speaker 0: Oh, I'll show you.
Speaker 7: So I'll show you, you can see just how many. Oh, yeah. Oh, yeah. Yep. Still messages.
Speaker 7: Still messages are just messages are still coming through on the queue right now. You can see there's a lot. See there's a lot of messages going through here. So, yeah, that's why. So, yeah, that's why.
Speaker 11: I have a potentially very stupid question, which is, in your capture the flag game. You said that we will not be able to get the flag TXT file, but then you ran the command that it just immediately became the flag TXT.
Speaker 7: Oh, that's because I'm the Yep. That's because I'm
Speaker 0: the owner.
Speaker 7: I'm allowed to have the flag. You're not allowed
Speaker 11: to have the flag. Right. How does it do that?
Speaker 7: That is done that is done using
Speaker 0: A sender.
Speaker 7: A sender tier. We essentially Right now the setup Right now the setup is not fully fully fleshed out, but it's just a base here sender tier that basically says I'm the owner. I'm the owner. Whatever I want. Whatever I want.
Speaker 7: You guys are you guys are Trusted rand untrusted randos. You can't do that. You can't do you can't my stuff. You can't have my stuff.
Speaker 11: So what about if, for example, somewhere in a GitHub that you asked it to go get? So if you asked it, hey. From this repository, in the remedial repository is x 3 but. You've maybe asked. You've maybe asked.
Speaker 11: It's going to go do it for you.
Speaker 0: It says
Speaker 1: you got your tier. Right?
Speaker 11: The thing you pointed at is inherently untrustworthy.
Speaker 0: Mhmm.
Speaker 11: Is that correct? It gave out of this model, but they would still have time.
Speaker 7: Well well Actually, no. Actually, no. In its current state, it cannot it can't fetch from a GitHub repository on its own, and the web fetches fetches are strictly white listed right now with the URLs. So essentially, we're building essentially, we're building this from we're chiseling this down from a secure design that's fairly functional and adding functionality incrementally. So right now, it's actually not cape not really capable of.
Speaker 11: But everyone you white list becomes a potential vector by which is somebody looted
Speaker 7: that's every site every site that you every site or address you whitelist will will oh 0 I'm glad you I'm glad you asked Kelvin Claw. Kelvin Claw is on GitHub at agent agent agent agentic highway. Highway. Mhmm. Oh, right.
Speaker 7: That'd be a good idea.
Speaker 0: Agentic
Speaker 7: Highway Agentic Highway slash slash Kelvin Claw on GitHub. The release The releases are out. Version version that I showed you is a development, is a development branch that I set up just for this demo. So but So but the version that's out right now is just released. The the I the Will the tall Arbash are tested on Linux and Windows right now.
Speaker 1: Should work
Speaker 7: should work on MacBook, but but that's up to up to someone to find out. And if it doesn't work, give give me an issue.
Speaker 0: Yeah.
Speaker 7: Yeah.
Speaker 11: Could that transfer the issue via sync to your club and the white list or divide?
Speaker 7: Can you repeat that louder?
Speaker 11: Same goal on which she by this white listing your Telegram ID. K. And you're gonna get it, but and then not having Not having your policy with any other ID.
Speaker 7: Precisely. Yeah. That's Yeah. That's, that's that's exactly what that accomplishes too. And And, that's How this how this would normally be operating.
Speaker 7: And the And the tier the idea of the tier users is that people can is it is it if you want other people to be able to, you can add them as maybe a trusted individual that's not necessarily, an owner that they can access some information that you care about. The and just like I said, this part's not fully fleshed out. In truth, we'll be probably 1 as more granular permissions and specific permissions.
Speaker 11: Why not even lock some track so you can just do that with on the cloud by creating a white bit IDs and black or black parts.
Speaker 7: I see. I see what you're saying. So this is Well, this was Well, this was all just for the sake of the demo. I wanted to see what would happen if Oh, no. Someone
Speaker 10: was able to
Speaker 7: The the someone was able to accidentally flip the flag or maybe the big man upstairs flipped a bit on my computer and opened it up somehow. I guess I guess
Speaker 0: what I
Speaker 7: was thinking here.
Speaker 4: Let's take 1 last question.
Speaker 5: Alright. Thank you so much.
Speaker 7: Thanks.
Speaker 11: Thank you.
Speaker 0: Yeah. Yeah. In the class. Come on. So do you want the last 1?
Speaker 0: Right?
Speaker 1: Yes. I think we are the last. Yeah. Yeah. So
Speaker 0: this is
Speaker 1: the last talk of the evening. And then after this 1, we're gonna do some lightning talks. So it's like a short round, where people can get up and just talk for, like, 30 seconds, say your name. And if you're looking for a job or if you wanna connect with people over a shared, interest, that's a good opportunity to, like, stand up. So just be thinking about that during this next talk of, like, if there's anything you wanna get up and and share.
Speaker 1: Sadly, there's not enough time for demos. Yeah. It's just Sun's Sun's computer. Okay.
Speaker 0: But before before we start, what do you want to build? Give me a few ideas. An app that tells me to cook. What? An app okay.
Speaker 0: Yeah. It does. Okay. Okay. Okay.
Speaker 0: I'm building the app. So the app is how to go. But next Will show you the part about, how to build that. I don't know what's the on this project, and then I will show you what our flow of Will. So it it Will take, like, 5 minutes or so.
Speaker 0: Right? Yeah. Okay.
Speaker 1: So 1 of the the cardinal sins of, AI Tinkers is showing slides.
Speaker 0: But I'm but I I
Speaker 1: do have, like, architectural slides that I wanna share with you guys. So so the premise of this is, like, in the future, what is it gonna be like when you can just tell a computer what application you want to build and it goes off and builds it and you come back and it's polished software. Right? And we're we're approaching that today. Right?
Speaker 1: We've seen, systems that can come pretty close, but inevitably, if you're building complex software, you're gonna run into some issues, right, with quality and that isn't what I meant. And, you know, the the AI will tell you that it works, but it really you know, it's lying. It tells you all the tests are passing, but they're not. So I did a deep dive over the last few weeks into software factories. And I'm gonna talk more about what I mean when I say software factory.
Speaker 1: And then we're gonna wrap up the talk kinda showing you how we've applied the software factory to the Sunday Club. We've essentially created a, what would you describe how would you describe it, Nikolai? Like, it's like a an AI hacker that comes
Speaker 0: Oh, I heard. Bill.
Speaker 1: Yeah. Okay. So, we're we're gonna talk about what is a dark software factory. I'm gonna talk about how you can define and refine specifications using something called OpenSpec and we're gonna actually gonna implement those specifications using another tool called Fabro. We're gonna walk through an example project with OpenSpec and Fabro, and then we're gonna build our first factory activated with OpenClaw and Telegram.
Speaker 1: And then we'll we'll wrap it up with some questions. Okay. So, just in the last year or so, we've seen a rapid evolution of AI development. And you guys may have seen this post by Dan Shapiro called the 5 levels from spicy auto complete to the software factory. So he's basically, like, chronicling the evolution of how we have started out with something like ChatGPT where it's very interactive.
Speaker 1: You know, you you ask a question, and it's basically like a smarter version of Stack Overflow. Right? You ask a question and the AI responds with an answer. Then we evolved to, like, level 1, which is where the AI is writing boilerplate code, an, unimportant code, but not like your critical code yet. Level 2 is more like an a junior developer, right, where you're doing, like, pair programming with this junior developer.
Speaker 1: And you're you're handing off some some control. The human is still very much in the loop reviewing that as it's generating code in real time. Then we get to level 3. This is where the AI is generating the majority of the code base and the developer is still reviewing everything that that the that the AI is doing. But in this situation, you may have found yourself in this situation.
Speaker 1: The human then becomes a bottleneck because the AI can write the software faster than you can review it. Right? And then you you realize you got all this code piling up that needs to be reviewed. So level 4, the AI can run for long periods unattended, you know, for hours potentially. And it can do much more complex stuff And then the human is is really trusting that, there's, like, self checks.
Speaker 1: Right? And you're only really kind of checking the final result. And then the final level, which is this is the 1 we're gonna talk about today is where the engineer is really kinda managing the goals, the IDENTIFICATION. It's defining yeah. Question.
Speaker 1: Where are we today? Well, let's take a poll. Who's who's who's in who's in spicy auto complete? Raise your hand. Who's in, coding intern level 1?
Speaker 1: Okay. We got a lot of advanced people here. How about how about junior developer level 2? Level 3? Okay, we have a lot of level threes.
Speaker 1: Level 4, senior developer. Okay. And how about level 5, soccer factory? Will, we got a few factories on this side of the house. Okay.
Speaker 1: Yeah. So different people are at different stages of this. Right? Because this is moving really fast. There's a lot of solutions out there and, it it may also be depending on what you're building.
Speaker 1: Your comfort level may if you're building, like, a a game, like, maybe you're okay with the factory. If you're building, like, health care software, maybe you're closer, you know, to this other level where you want more control and oversight in it. Okay. So there's where where I heard about dark factories from Simon Wilson. Who reads the Simon Wilson blog?
Speaker 1: Yeah. It's it's a great blog. And he talked about this, tour that he got of strong DM and their premise is that all code must be written by the AI. So, like, no code can be written by humans and not only that, but all the code must be written Will the code must be reviewed by AI. Right?
Speaker 1: And okay. Think about that for a moment. All the code is written by AI. No humans involved, and all the code is reviewed by. And so in practical terms, this is like, okay.
Speaker 1: If you haven't spent 1000 dollars on tokens today, then your software factory has room for improvement. That's kinda like the the idea. So the question that Simon poses, which I I'll ask you all is, like, how could it be pass how how could this be possible, possibly a sensible strategy when we all know how prone LLMs are to making mistakes? Right? We've all we've all faced app’s situation, right, where the LLM says that it's done something and then you look at it and it's like it's completely failed.
Speaker 11: So that's I think this
Speaker 1: is kinda like the holy grail. Right? Everyone is trying to figure out, like, well, how can you get the AI to actually when it says that it's done something that it actually has done something. So we've all heard of technical debt. Right?
Speaker 1: If you've been in the engineering space any amount of time, you probably have experienced, you know, debt over time. You have, you know, code that goes stale, stuff that changes, but it's still in the code base. Well, now with AI, it's really accelerating the pace of not just technical debt, but cognitive debt. So the engineering team knowing what has been Will, but more importantly, the intent debt, which is why has this been developed? Why did this change?
Speaker 1: Right? And and this isn't always getting captured when you're moving very fast with AI. A lot of these decisions are getting lost. And then you looking at the code base, like, why did it do that instead of that? So 1 of the tools that I've been exploring to try to address this is something called OpenSpec.
Speaker 1: And the best way to describe it is like a lightweight spec driven framework. And Will, the idea with this is that it captures the change in requirements of the system. So this makes it a lot easier for developers to understand how they're modifying the system and what still needs to change. So you can review the spec and you can see, oh, I understand now why we change the code from this to this. Right?
Speaker 1: And this is, like, really important when you're moving at AI coding speed because there can be a lot of changes happening. And if you just let the AI kind of run amok, it can make a lot of changes and then suddenly your your code becomes unrecognizable. So OpenSpec is a way to kind of capture all those requirements as you're going and snapshot them and have them in your Git repository. And that also becomes a good resource not only for humans reviewing the code base but also for the AI. It can kind of use that as ground very grounding mechanisms or like source of truth.
Speaker 1: It can look at those specifications and know, oh, I can see why we made that change. Right? Okay. So let let me talk about Sunday. Nikolai, feel free to chime in here.
Speaker 1: Did you get a microphone, by the way? Oh, you didn't get 1? What? Oh, here it is. Yeah.
Speaker 1: If you Wafaa put this on. I don't know why it's flashing. I'm a beginner. Okay. So we took this idea of, like, OpenStack and something called Fabro.
Speaker 1: I'm gonna talk about in
Speaker 0: a sec. We thought, what
Speaker 1: if we could just tell OpenClaw via Telegram what we wanted to build and then have an entire pipeline, have the AI checking and validating and doing a lot of the steps that normally humans would be doing like making the GitHub repo and, you know, deploying the software up to Google Cloud. And so this is a 15 step process. I'm not gonna go through all these things, but you can see there's a lot of stuff happening here. And at the end of this pipeline, what comes out is the GitHub URL. So it's it pushes the code to GitHub, the URL on the project page on the sunday.club website, and a deployment URL.
Speaker 1: So where the code has been deployed where you can actually try it out. And then there's some error some AI verification steps as well to make sure that the the soft
Speaker 0: tell us what SonicCloud is. SonicCloud is a hacker attack every time. So our goal here was basically to replace our call, with AI. So AI can hack, and do everything we do every Sunday, by June. But question?
Speaker 0: Still the number 9 there. I
Speaker 1: Like Well,
Speaker 0: there is a
Speaker 3: side of the website with our projects.
Speaker 0: Each project has likes. So both can like its own project, so it has a bit 1 like. It's like liking your cost.
Speaker 1: It's like it's like the tip jar. You know? If if you're a musician, you put, like, a dollar in your tip jar. Let's see if we actually have I had it in here somewhere. Okay.
Speaker 1: I have too many tabs open as you can see. Okay. Back to the just have, like, 1 1 or 2 more slides here. So number number 2 in this why is it saying loading now? Okay.
Speaker 1: So number 2, which is create GitHub repo and open spec Fabro. This is a this is zooming in on just number 2. Okay? So what Nikolai did at the very beginning of the talk is he asked you guys what you wanted to build and he typed that into Telegram. Then what happened is that message got sent to an open cloud gateway and on which we have a skill called Sunday project pipeline.
Speaker 1: That's step 3. And that's that 15 step orchestrator that does all this other stuff. Right? That Sunday project pipeline calls the open spec workflow, which essentially distills a full spec from from that initial query. So, like, you just give it a 1 sentence description and it goes and flushes out, like, the design dot m d and a task dot m d and a proposal dot m d and, like, a whole bunch of other and for each feature in your application, it builds out a specification.
Speaker 1: Then that entire spec directory gets passed to to Fabro, and we have a special workflow called Sunday ship. And what Fabro does is it takes that task dot m d file and it breaks it into a graph. And in that graph, we have conditionals, we have loops, we have, quality gates, verification checks. So by the time that software gets deployed up to Google Cloud, it's gone through a very rigorous process of, like, checking the quality, linting, doing a lot of, like, you know, looking at the plan and making 2 versions of it and then synthesizing the best of both both plan. There's, like, a lot of machinery happening.
Speaker 1: You can see these are all the files that are getting generated behind the scene. So I'm gonna hop over and show you guys what's actually going on in the code base. So this is our hacker repo. This is up on GitHub if you wanna check it out. That's the URL, sundayclaw/hacker.
Speaker 1: Yeah. So at the beginning of the at the beginning of the thing, I I put in the same thing that Nikolai put into OpenClaw. I just did it on my computer so you can see what's happening. OpenClaw can be a bit of a black box, right, when you you tell OpenClaw to do something. This is part I've been little frustrated.
Speaker 1: Nikolai, I think, has gotten a lot more comfortable just telling OpenCloud to do something and trusting that it's gonna do it okay. I I like to see what's going on. So I have it running on my laptop. So we have these things running in parallel. He's running it up on the cloud, and I'm running it locally so you can kinda see what's happening.
Speaker 1: So here's the learning to cook. So the this was where we put in the prompt. Can everyone see this bio, or should I make it a little bit bigger?
Speaker 0: Okay.
Speaker 1: Alright. So this is this is basically the same thing that's happening up on our OpenClaw server. It's just we're replicating locally. Not right now. Right right now, we've we've auto app’s.
Speaker 1: We do have some human in the loop steps where it'll it'll pause and ask you, is this is this specification look correct? We've until we can figure out how to make that interaction on OpenCloud, like, smooth, we've just and also for demoing, it's a little better to just have everything happening automatically. But we do wanna introduce those human loops so that once it does the first version of the spec, it'll pause and, like, ask you to review it before it continues. Okay. So yep.
Speaker 1: It can run-in parallel, but I don't know right now. We haven't really tested this at, massive thundering herd.
Speaker 0: Yeah. Yeah. Yeah. Practically speaking, like, when we tried it yesterday, yeah, it just, like, waited for while it's building Nate's project, it waited before it starts to build my project. Yeah.
Speaker 0: Should I show the
Speaker 1: just just 1 0, is it is it finished already? Yep. Okay. Beef before he does that, I just wanna show this is the workflow kind of what's happening. This is a simple a simpler version of the workflow that I'm using locally.
Speaker 1: This 1 is the 1 on this on the Sunday Club website where you basically do do a planning phase, then there's an implementation phase. Then we install the dependencies. We make sure all the dependencies install correctly. We run we run a lint check. We check to see the lint pass.
Speaker 1: We build it and then there's a there's a loop. So it goes if there's any lint errors, it it loops back and fixes them. And if the build fails for any reason, then it goes back and it kind of, self improves. Basically, loops back on itself and then it'll just keep debugging until it until it gets everything to pass. Do you wanna show some
Speaker 0: Yep. Wait. Wait.
Speaker 1: Do I need to admit you? Wafaa, do you know if she's No.
Speaker 0: I'm sharing. Oh, you are sharing?
Speaker 1: Oh, do maybe I need to stop sharing. Maybe.
Speaker 0: Okay. Try now. And you. Yeah. So, actually, I skipped Fabra, because it takes, like, 10, 15 minutes, and we don't have this time.
Speaker 0: So it it just, like, tried to Will, an app’s, with with AI model itself. So this is a Sunday card. Actually, it experienced some issues with Sunday website, so it's not like, Sunday card is not fully working. No. We put
Speaker 1: a full description there.
Speaker 0: Yeah. It it it missed the description and link. But, it Will a so I asked it, build a new Sandy project, give Fabro, and that tells me how to book. Wait. So someone's idea.
Speaker 0: And it started building, it Will a you app’s repo, for this app. It it drops back with, open spec. Then it straggled a little bit with some, some coding issues. Then it does a verified, deployed the GCP. So this is our cook compass.
Speaker 0: And, helps start the near what what does it do? I don't know. You you picked the recipe, it looks like. Yeah. Oh, I just have
Speaker 1: I have egg. Oh, you can say what ingredients you have.
Speaker 0: Yeah. We Will see. So it Will an app, basically, that users AI. It use the 3 d open router endpoint here, so we are not running out of credits, but it's kinda stupid. And what what what what what's going on?
Speaker 0: Who understands this interesting UI?
Speaker 1: There IDENTIFICATION.
Speaker 11: Wait. I I pressed find my
Speaker 1: This is what happens when you give a 1 line specification. Yeah. It it just kinda goes and builds random stuff.
Speaker 0: Yeah. Probably, like, l also if you have to use the Fabro, it would be
Speaker 1: Yeah. If they did use Fabro, we we we better we probably would've gotten a better
Speaker 0: Yeah. But Okay. Oh, we don't know. But I used our friend's, openly idea.
Speaker 1: Yeah.
Speaker 0: There's lots of credit. So I hope not much. But not at all. This Sandeepro, it's a pipeline. It's a skill for OpenClaw, that creates Sandeeprojects, from GitHub report, like, deploying, create a Sandeep page, based on human request or randomly.
Speaker 0: So I I usually, also create some random projects. They created the AI Tinkers website, and they've hallucinated, our schedule for the day. Kinda looks similar right here. And, yes, other projects.
Speaker 1: I was gonna show the the skill that that actually does the work. If you don't mind,
Speaker 0: stop and share. Yeah.
Speaker 1: Okay. So so this is the actual skill dot m b file. So you could see it's quite involved in terms of, like, instructions and, you know, step 1, phase 1, phase 2, phase 3. And and and the claw the open claw agent is basically following this, not religiously, but it's it's following it as close as as as it needs to to deliver the final thing. And we found that it it it's pretty smart about it.
Speaker 1: Like, if it reaches a block, it'll kinda find another way to do it, which sometimes means skipping some steps, which, so so there's still, you know, there's still some room for improvement here. But this I think didn't you say that the hardest part was just getting it to be able to post to this like, it wasn't building the app. It was getting it to post to the project page. That was, like, the most complicated piece.
Speaker 0: The site website doesn't have API, so it tries to click buttons. And, it's, it's harder than building a software, for AI.
Speaker 1: Yeah.
Speaker 0: Yeah. And then you have those that we were joined earlier. Mhmm. That that was a I think it's fabric. Yep.
Speaker 0: Yep. Yeah. So what like, actually, 1 step out of this system, use cover it. Yeah. Yeah.
Speaker 0: Yeah. But that that Right. Logos. Yeah. Yeah.
Speaker 0: Yeah. Yeah. Right. That's that's what we find.
Speaker 1: Yeah. So the the these are the artifacts that get created. Actually, that's not the cooking 1. Let me get to the
Speaker 0: Maybe we can.
Speaker 1: Yeah. Will I while I'm pulling stuff up, we can take some more questions. Sometimes.
Speaker 0: Sometimes. Not always. Okay. Sometimes. But, yeah, it did app’s wrong thing around this step.
Speaker 1: Yeah. So so there's obviously, like, tests. If if you look at this, this is like So so this is this Sergeant this is another workflow that I asked, Claude code to build. So Fabro has a skill called create workflow. So when you have, like, a code base, you can tell Fabro to go create a workflow specific to any particular thing that you wanna automate.
Speaker 1: And so I I told Cloud Code, use the Fabro create workflow skill to create a workflow for doing browser testing. Because I found it was very annoying to have to, like, open up the browser and click on stuff and then tell AI that, you know, something was not working. So this is a workflow that it came up with for a game that basically, like, runs various test suites, you know, visual visual verification, gameplay control, scoring progression. It collects all the results, and then it checks to all the test pass. And if they do it, it generates a pass report.
Speaker 1: If they fail, it generates a fail report, and then it exits. So this this is an example, Andy, of, like, how how you can provide some, like, end to end browser testing that once you give the AI access to your logs of the front end and and back end server and you give it access to a browser, it can sort of self heal. It can, like, figure out, oh, I can see in the logs here. I'm getting this error, and then it can see in the browser where it's things are breaking, and it just iterates until it fixes it.
Speaker 5: That's not the solution that this
Speaker 0: No. Not not yet. Not necessarily. Some of the steps like a deployment, then adding links, like, publishing on Sunday club website, adding links or to GitHub to Sunday project. Those some of them, like, logistically, capturing in building.
Speaker 0: Yeah. Yeah. It it it does. Yeah. Yeah.
Speaker 0: It it said it does the and, yeah. Sometimes it takes screenshots and, like, find the that's whether, like, still something works well and, look good and so on. Yeah.
Speaker 2: There is something that I said, what I told you, I had the same issue with who called it out of, out of order or something will not happen. Yep. I mean, what tells me what I did was using something called faster output And then having a code that does the steps there. So that's, it goes to x y z and no Yeah. There it
Speaker 0: Yeah. For us, what helps, is very clear. Repeat, like, steps multiple times. And, also, we have app’s skill of 15 steps. And, also, there's separate checklist.
Speaker 0: The checklist that basically have most of the app’s. And last step on a skill is to use checklist and verify that everything is done. So I did, like, multiple parts of the skill asks AI to stay on track to everything and check that every every book. Cheaper model doesn't do Will, with open flow in general, and and with, like, 1. So it's like, for example, I think minimax was really important.
Speaker 0: So it's the GPU 5.4. But yeah. Like, 3 model, 3, other model, like, you know, do it almost like anything. Like, Minimax works sometimes. No.
Speaker 0: It's more like model coding and overall intelligence model. It's not a it's not a complex type. Because it's it's a multiple, call like this. Like, it's kinda like do sub prod automation. But overall intelligent, and.
Speaker 2: And my question is about the back preference.
Speaker 0: Have to make sure you're correct, verify yourself, something like that. So it was like pretty.
Speaker 1: Yeah. I I wanted to mention that Fabro has a kind of experimental web InterSystems is all dummy data right now, but if you go to the GitHub repo, you can can try this out. So this this is basically taking all the command line artifacts that are created and try to visualize them. So these are all my sessions that I have, like, just chat sessions with the app’s. That 1 is not working.
Speaker 1: These are the different workflows. So each of these is represented by a diagram. So you can see, you know, kind of how it diagnose the root cause. It applies a fix. It validates, reviews the changes.
Speaker 1: Implement feature has a has a different workflow graph. Well, actually, it's the same. And then these are different verification steps that you can
Speaker 0: specify. It's hard to measure. How do you measure? How do you measure what? Face.
Speaker 0: Yes. Yes. So we have ideation phase, like, the first first 1, and you can also take that idea randomly. And it what it does, it takes best, like, Sunday projects and rely like, have some inspiration on already existing projects on app’s Sunday club. So since we, as a as a hacker club, like that project, maybe, like, we Will like new 1.
Speaker 0: But, overall, it it's, like, it's not very sophisticated idea.
Speaker 11: Yeah. Some of
Speaker 1: the improvements are, like, adding UI UX skills. So there there's some Yeah. Some there's this 1 called, impeccable
Speaker 0: Mhmm.
Speaker 1: Which is, like, design fluency for AI harnesses is how they describe it. So it has, like, 20 different design skills. And if you add this into your cloud code skills directory, which we can do Will Sunday cloud as well, suddenly, you've just it's like, I know kung fu. Now suddenly, your your coding assistant now knows how to write how to make designs, and they don't all look the same as like, you notice that the the last few the apps that we built all kinda look the same. So that this gives you you know, it gives your AI, like, better design aesthetics and
Speaker 0: In the
Speaker 1: middle, I think,
Speaker 0: since underlying models and tools are improving, in a year, the same ThunderClaw, like, even without any change, it could build much better tasteful product. Did we
Speaker 2: come with anything?
Speaker 1: How to capture what? Where yeah. So so it is creating a, an instant log. So as it's going through the workflow, whenever it encounters an issue, it logs it to markdown file. And then you can feed that markdown file.
Speaker 1: Well, also JSON. It has, like, JSON l files. So you can feed that diagnosis report back in to the build or the implementation phase. So it'll it'll just keep iterating until it's resolved all the issues in the list. I think we may take 1 or 2 more questions.
Speaker 2: So, the fibro, I think, is, like, deterministic graph behind it that that that controls its flow and your kind of order 15 step process will come up
Speaker 0: more like, off the 11 driven workflow. I mean, that is
Speaker 2: what your experience is in those and, like, where you see the strength weaknesses. Like you said, it might
Speaker 3: it might self heal better, but then
Speaker 2: it gets done. Like, yeah. Where where do you where do you see that balance? Where do you see that balance? Where do you see it from?
Speaker 0: I think we're kinda different to make, well, I I I I I did on the LLM side, and just give it Will, give it looks, and then they'll figure it out Yeah. Somehow.
Speaker 1: I'm a little bit more on the deterministic side. He's he's more on the just throw it over the wall and let the AI figure it out.
Speaker 0: Yeah. So I think our first software development, Avro, like, works great for less deterministic things that can eat less deterministic function. Yeah. But there is no silver bullet here.
Speaker 1: Yeah. So these are some examples of, like, the logs that get generated when you're running, Fabro. So it's it's showing you sort of it's checking all these different things like correctness of the different API endpoints, security, doing a bunch of security checks, simplicity, completeness, consistency, and then it it documents any changes that it makes as well. And this is example of, like, an implementation log where it tells you what files it created, the source files, the front end, the test that it created, any issues that it encountered. So
Speaker 0: I'd love
Speaker 3: needing or iterating on,
Speaker 11: you know,
Speaker 0: vector than several adult application. What's that like? What what are
Speaker 3: the systems that we're launching something? Just wanna change more component. Is spec is a
Speaker 0: a component in the code or is it
Speaker 1: Yeah. So so the way OpenSpec works is is you type in, o p s x o p propose. If you wanna, like, brainstorm, you do the explore, and then that kinda launches into, like, a Socratic thing where it'll start checking your assumptions and asking you questions and try to try to refine and figure out the shape of what you wanna build. You really know what it is you wanna build, you just do app’s. And you just, you know, tell it tell it what it is you wanna propose or what you wanna build.
Speaker 1: So I would say something like,
Speaker 11: add a
Speaker 1: cooking timer feature. Right? I think because it's doing this up here, it might might not also let me propose a feature at the same time. And then those those new specifications will get defined over here in this changes folder. So you say I have changes, and this was the original specification right here.
Speaker 1: Right? So this this is was the proposal. My screen is kinda tiny. Sorry. I can't show you that right now.
Speaker 1: But but, basically, as you define new features, they'll show up here with the timestamp. And then when you tell it to go build it, it knows what it's already Will, and it'll just kinda pick up where it left off. So you can see it's
Speaker 0: change app’s back and open. Give you
Speaker 1: Yeah. And and because all the specifications are stored in your git repository, you can also see how how things are changing over time. You have full version control of your specifications.
Speaker 3: But OpenSpec isn't doing anything. It's just writing
Speaker 0: We're gonna writing
Speaker 3: writing that's writing. Yep.
Speaker 0: Is it right and true way of
Speaker 1: Yeah. So so here you can see, like, when I did that OPX app’s, it's asking me, are you doing a new feature, building something from scratch? Are you doing an enhancement? Are you doing a bug fix? Like, what are you doing?
Speaker 1: So I can say, like, I'm building a new feature. And now it's gonna start writing out the specifications for that.
Speaker 0: K. Thank you. Thank you all. Okay.
Speaker 1: So at at this point, we're gonna move into the lightning talk round. So if you would like to come up in front of the group and say a few words, like, 30 seconds to a minute, it could be you're looking for a job. It could be you're hiring. It could be you're looking for other people to work with you on a project. Yeah?
Speaker 1: Correct.
Speaker 2: Okay.
Speaker 0: Slash try.
Speaker 1: Iteration slash try.
Speaker 10: Okay.
Speaker 1: Okay. So if if anyone wants to come up and talk, just come up to the front and just we can start lining up over here. Just say your name and what you're what you're looking for.
Speaker 0: Sure thing. K. My
Speaker 3: name is Brian Hirsch.
Speaker 2: I just Wafaa make
Speaker 3: a quick plug. So a project that I've got really into or runs behind the product I'm developing is called Eads. So some people have read about Gas Town or Gap City, P I E. So, these is awesome. I've been using it since December and, horrific.
Speaker 3: You you are a skeptic about, like, can you actually buy something that, like, works and it's Will and usable? You know, these has, like, totally Anyways, I have a little side project, beads. I've contributed a couple of small patches. I submitted a proposal that beads is into about, getting beads to track DOLT history in a way that corresponds better with the way we track Git history. So DOLT, which is the back end of feeds, is like Git for MySQL or Git for a database.
Speaker 3: So right now, the 2 histories diverge. It's kind of a it's like, super intuitive because it sort of feels like Jira when you start using it. Then we start thinking about it. It doesn't actually make any sense. So, anyways, I got the proposal in.
Speaker 3: I started doing some testing
Speaker 0: with it, and I'd love to find some people to collaborate Will, whether it's
Speaker 3: a sounding board or testers or, like, 1 development stuff together. If that interests you, come
Speaker 11: chat with
Speaker 3: me. Thanks.
Speaker 1: Yeah. I forgot to mention when I was, showing off Fabro that Beat was another 1 that I I looked at. We can't wait to They have a lot of similarities.
Speaker 0: I asked about that after.
Speaker 3: Hi. My name is Jeff Levinson.
Speaker 0: I actually run the AI Takers of New Hampshire, and I've, been coming to Massachusetts AI events here so much. We're brave in App’s North. And we have also AI weekly Hampshire Avenue in the week of, April starting April 13 through the seventeenth that week. Sponsoring a free AI training workshop and was interested in that too. So you can find who you are on LinkedIn, you know.
Speaker 0: So officially What
Speaker 1: what's your AI concurred?
Speaker 0: AI in New Hampshire. It's Manchester. It's Manchester, New Hampshire AI takers.
Speaker 1: Alright.
Speaker 0: So that if you Google Manchester, New Hampshire takers, then that's the 1 that we run, and, Travis presents there quite often.
Speaker 1: What are the dates of the event?
Speaker 0: The next one's gonna be the fifteenth, and that's in conjunction with AI Week New Hampshire. So anyone who wants to check out that website, we can go. Alright. Aaianewhamster,uh,.org. Alright.
Speaker 0: Let's go on. Thank you. Alright. Thanks, everyone. Yep.
Speaker 0: Yeah. Hi. My name is Robert. I'm a software engineer. I work at a
Speaker 3: company called Freshworks. I sold my company FireEye here
Speaker 0: for them about 3 months ago.
Speaker 3: Don't Will my new employer doing this. I have a side project, that's, agent orchestrator. I'm just curious how people are solving this problem. So it's agent orchestrator in the cloud called Cadenia. If you go to that.com right now, you'll see another fire, but I promise it's real.
Speaker 3: I'm just curious if anyone is doing anything in agent orchestration. I'm curious
Speaker 0: what it wants to do. What's the URL again?
Speaker 3: Cadenia.com, but it will not go anywhere. But I know How do
Speaker 0: you spell that?
Speaker 3: Cad well, actually,
Speaker 7: how do you spell
Speaker 0: it? Cadenia, c a d e n I a. I know
Speaker 3: I gotta buy that 1. Oh,yya.com.
Speaker 0: Tedenya.com.dotcom.
Speaker 3: On the k version too. Okay. But that's the site. That's what it will be. If you go to hap.kadenia.com, you can sign up and use my credits.
Speaker 3: Yeah. It's up to you.
Speaker 0: And what is it? So it's an
Speaker 3: agent or orchestrator. I I'm on the train of agents who Arbash
Speaker 1: only as good as
Speaker 3: Will give them. And my belief is that if you work in an enterprise or even Will company, you likely have APIs for your application already, rest to your PC, MCP. And, Cadenia can talk to any of those transports, and you can actually create, basically, an agentic harness, agentic loop, comes with human in the loop app’s, and everything is API based. So it comes with TypeScript SDK. You can ask, then, you know, what is the agent doing?
Speaker 3: And I can also send you webhooks for when it asks for approval. They'll send you webhooks for any event, any tool call that app’s actually performing.
Speaker 8: And, yeah, Will little side project.
Speaker 3: But if you sign up, it'll give you a little tour now.
Speaker 1: It's nice.
Speaker 8: Alright. Hi, everyone. I'm Seth Weitzman. I work as a software engineer, and I work for a company called Enfy, e n f I, bring that in there, dot a I. We do agentic loan processing.
Speaker 8: So 1 of the problems that a lot of banks are having is they have a lot of money they wanna give out. They don't have enough people to actually approve those loans and do the underwriting. We are working to make that a lot faster. E n f I dot a There you go. And, that's kind of fun and cool, but what they're also working on or we're working on a lot, we have a lot of smart people, which I'm learning from, is we're doing a lot of the code factory stuff.
Speaker 8: So we're doing a lot of the tooling and creating of these things that can kind of build this stuff on its own or faster, remote development. We just tell our our our, sorry, our our, project managers can do triage because they can just tell Slack to go fix it. They don't have to have a development environment to work and really, increasing the speed and, efficiency of all of our of our projects. It's really cool. I'm new to AI.
Speaker 8: I just this is my first AI job. I've been software developer for 20 years. I'm putting in MRs, daily in a language I don't know how to write it. I don't know go, but all it does. And so it's a it's a really cool and very fast paced environment.
Speaker 8: 7 engineers right now. We're looking for a principal, a senior, and. So if you wanna check it out, they're on the website or feel free to grab me. Thank you.
Speaker 0: My name is Emea from Origami Grants. We're a new startup focused on making AI development workspace for, grant proposal development. I'm asking for if anyone has experience with using LMS on PDF data extraction. I would love to learn from you on things that went well, things that didn't. So on network.
Speaker 1: What's the name of your organization?
Speaker 0: Orgamigrants.com. Orgami. Orgami. Orgami. Orgami.
Speaker 0: Orgami.com.
Speaker 12: Hi. My name is Prachi, and I'm doing the start up value of jigsaw.com. And, basically, our goal is to help people understand what the AI is like a look. How keep track of a software that's been developed. Yes.
Speaker 12: So we just GA'd last week, just a week before that. And so we would love for you guys to just try it out. You can have a local installer or a platform. Basically provide the code repositories and we create a software architectural diagram for you so that you know what the software looks like. And we provide some additional things if AI did some making big code changes and visualize it.
Speaker 12: Actually, how are they making the changes? Where are they making the changes? You can dive deeper into those things. Now, obviously, we want to do all of it even more. But, would love for you guys to try it out and give feedback.
Speaker 12: Good, bad, or anything. And, yeah, we also open source. We Arbash actually launching bunch of pages, which are, architectural these open source software libraries. So if you're using multiple clause, we have created an architect for Nanoclon, NemoCloud. We will be doing some more, open source repository.
Speaker 12: So if there's any particular repository that is scary and you don't know what's happening inside of it, we can get an architecture diagram for you guys and, you know, make it available for the world. Thanks.
Speaker 1: Great. Alright. Anyone else? Last Will? Okay.
Speaker 1: Well, thank you all for coming. As as always, if you wanna give a talk, can talk to me or any of the of your other organizers or if you Wafaa sponsor, we're always looking for sponsors. And I think we've got the room for another hour or so. You're welcome to stick around, have some more pizza, take 1 home if you want, and we'll see you next month. Thank you.
Speaker 1: Thanks.
Speaker 0: Hey. Yeah. Super inspired by the, Android.
Links
Tech stack
Finding related talks...
Compose Email
Sending...
Email preview
Loading recent emails...