Defending Open Source in the Age of AI

glyph bracket orange
glyph bracket orange

AI is changing how software is built, secured, and attacked. Used well, it can help experienced engineers move faster and uncover issues earlier. Used without the right context or review, it can flood maintainers with low-quality findings, introduce insecure code, and erode trust across the open source ecosystem.

In this episode of Open Source Open Mic, Sonatype’s Andrew Garrett speaks with Christopher “CRob” Robinson of the Linux Foundation about what responsible AI adoption looks like for open source. They explore how AI-generated vulnerability reports are affecting maintainers, why high-quality disclosure and remediation matter, and how teams can prepare for a future where AI accelerates both software delivery and security risk.

Transcript

0:05 Welcome everyone to this episode of Open Source Open Mic. My name is Andrew Garrett and I work in product marketing here at Sonatype. I am pleased to be joined today with my guest Christopher Robinson aka Crob from the Linux Foundation. And Crob, do you want to do a little introduction for our viewers today?

0:25 Absolutely, Andrew. Uh, thanks for having me today. My name is Chris Robinson, also known as Crob. I get to do security stuff on the internet and predominantly what that means is I am the CTO for the Open Source Security Foundation. I am the CTO for the new Accrees Foundation and I'm the chief security architect for the LF. So whenever there's security things, I get brought in to be helpful.

0:55 Awesome. Yeah. And Crob, I'm really happy that you carved out a few minutes of your busy schedule to speak with me today. I know you've got a lot going on. You were at Black Hat recently.

1:05 Oh, I know. It was so hot.

1:07 Yeah, I think you said 115 degrees.

1:10 You were 115. I thought I was going to die. Man, well, again, I don't know why they decided to have Black Hat in Las Vegas in August, but that's a topic for another podcast.

1:21 It's crazy. Yeah, Hacker Summer Camp is always fun. The whole Project Diana, Bsides, Black Hat, Defcon. I got to do most of it this year and it's always fun catching up with old industry acquaintances and having interesting security conversations.

1:38 Yeah, absolutely. Well, our security conversation for today is going to be centered around defending open source in the age of AI. We've been hearing a lot about AI this year. I mean, software development is changing fundamentally with the onset of AI. Any thoughts before we dive in? Any thoughts on just kind of the AI momentum we've seen over the last 12 months or so?

2:01 Yeah, it's interesting. My gray beard cyber friends and I were at a conference last year and talking about it and one of my friends actually has his master's degree in machine intelligence, which is what all this nonsense used to be called. and he got this degree like 20 years ago. Uh and it's interesting he did a whole talk talking about the different AI techniques and at the end of the day all these collectively these tools are math that it's an algorithm trying to predict what the next thing next word would be or the next most likely piece of data that needs to be served up. But it's wild.

2:49 We collectively the ecosystem really ran into a speed bump last fall is when we started to at least in the open SSF community. We were talking about um this term called AI slop where people were leveraging things like chat GPT or claude or one of the open models to scan repos and then they would just f these researchers would just throw a report over the wall at the maintainer and at last fall those reports were super high volume very low quality and really wasted a ton of time on maintainers and left a really bad taste in open source developers mouths and it's that's part of what I get to grapple with today.

3:32 Earlier this year, probably the February time frame, something changed in the frontier models especially and then cascaded down to the public and then the open and uh local weight models where they got pretty good.

3:51 The accuracy and the fidelity of the findings were better. they they still are super verbose and sometimes they still hallucinate but they got better and that's where you know our topic of conversation today is going to kind of dive into some of now that we're in a space where AI is much more usable you know how do we kind of earn back some of the trust of the ecosystem and then you know encourage everybody to start using these tools because these things move at a velocity that we have not seen in 50 years of software engineering. It's crazy.

4:25 I'm glad you brought that up because Yeah. In software engineering in particular, the theme has always been speed, right? You know, we want to move as fast as we can. Uh we don't want security to slow us down.

4:40 And so it seems like there's an opportunity here for software developers to really, you know, if we can use AI the right way, we're going to be able to develop at speeds that used to just be a dream and now it's actually a reality.

4:57 Exactly. And that's and it's interesting you all Sonotype did a report um it's been a couple months back and it's a the AI is kind of a double-edged sword in the hands of an experienced engineer you know somebody that understands the discipline of creating software from an engineering perspective this the force enabler on this stuff is ridiculous like you can legitimately get four to 10 times your work output if you're leveraging these things as assistance.

5:25 But from a junior perspective, what we've noticed is a lot of people that aren't traditional software engineers are now writing software they're vibe coding and the results of this stuff the quality is very low and if you don't have that knowledge you don't have that training and experience to understand and you're just blindly trusting what the robot gives you we're going to have some bad days.

5:50 That's why I always get a little like a little afraid when I hear people talking about oh I vibe coded this app in an afternoon and I'm going to replace insert ridiculously large commercial package with my vibecoded thing. Like, yeah, sure you are. Good luck. Uh, Godspeed.

6:07 Uh, we'll probably be talking to you in a couple weeks after it, you know, falls down.

6:12 So, Crob, are you telling me that with great power comes great responsibility?

6:16 I am.

6:17 Was Uncle Ben right all along?

6:19 Uncle Ben was absolutely right. Again these tools and I come at this from my perspective we only have three problems with AI.

6:31 Using AI to write secure software which is kind of this topic right now using AI to secure software which is what we'll talk about with Akrites and then finally it's the security of AI systems so like if you are paying attention.

6:47 If you've read the news in the last couple weeks the frontier model folks have had some oopsies and the robot escaped and that's again I would throw that into that last category of these systems weren't effectively secured and the most plausible event actually happened which you know the thing broke out you trying to follow its own instructions.

7:10 Yeah you know I find that we are living the plot of iRobot more and more. It's no longer fiction, we're living it.

7:18 Right well I like to think that I'm Kyle Reese helping humanity fight the robots every day.

7:26 Absolutely. You brought up Akrites. I do want to dive into it. Uh backstory. I did look up the word Akrites because I was curious where this name comes from. Um and apparently the term Akrites refers to the legendary frontier guards and border militias of the medieval Byzantine Empire.

7:48 Yes. So I need to ask you first of all, did you come up with this name?

7:52 No. Okay.

7:53 No, no. We we had a whole committee where we had I wanted to call it robot fights and uh my vote lost. But yeah, Akrites is really interesting and apropos and that's one of the things I love about open source is we have like the most there's so much creativity and like the name is so evocative like Kubernetes which is an old Greek word.

8:18 Uh so the Byzantine Empire, these frontier guards, they were out on the periphery of the empire and they would were watchful for invaders and then you know as they saw like an invading armies they would go they would part of them would help defend and the other parts would go back and help warn uh the home the home guard and that's really where we are like thinking about where we are in the space with software development and leveraging these large language models.

8:44 We really are out in the frontier, out kind of out in the wilds, and we're there to try to help protect both upstream maintainers and projects, but also the whole downstream community that there's a lot of good things that can happen using these tools if they're used correctly.

9:02 But there's also a lot of bad things that can happen. Whether you're shooting yourself in the foot because you didn't use it right, you didn't read the instructions or the bad guys got a hold of it and they read the manual and they're doing things uh they're doing naughtiness towards you.

9:18 But that's what, you know, again, we're there to try to help and try to be that front line of people and help cushion the blow for the oncoming wave that's going to hit us all.

9:26 Yeah. And I mean the Linux foundation uh is really centered around security fixes for open source. Um so let's talk about Project Akrites specifically.

9:41 What is it exactly? You know for starters let's set the stage. What is Akrites? And what are the main goals?

9:48 Yeah. So u has a couple main functions. Predominantly, we are here to help the whole ecosystem deal with the use of LLMs and other tools that are creating high volumes of security reports. And Acci's mission is to take those reports tighten them up, making sure they're of high quality and they have all the things that an upstream maintainer needs to make decisions on that report, and then get those things safely disclosed upstream so that the upstream maintainer will take that work either take a patch we give them or write their own patch and then the whole ecosystem benefits.

10:35 So we're kind of again that we're the vanguard out front trying to help make sure that if we get upstream prepared all that work and patchwork will filter down into commercial enterprises and then down into you know banks and retailers and restaurants and whatnot.

10:53 And so we do this through the creation of an upstream uh incident response team. So if uh depending on where your listener lives in the world, they probably have a group called a corporate security incident response team. So if you're inside a company, you have the information security cyber security people. Um but if you work more in a product role, you probably are familiar with a product security incident response team.

11:17 Those are the people that make sure that the goods and services that are sold to customers are of high quality. So like you know Sonotype has a security team that pays attention to all the different offerings you all have.

11:29 You know, we're trying to be, you know, we aren't product people and we're not an inside corporate entity, but we're trying to perform that incident response function where people will report things to us and we will make sure that they go they land in the right spot and then when things go public, we make noise and make sure that everybody understands on this day this problem was fixed and here's how to get the patches.

11:53 Yeah, that's really been something that's lacking in the industry where I have personally been talking about doing this for about a decade. We've had plans on the books at the OpenSSF for about five years and it's just a matter of there was never enough uh resolve or grit to commit to doing this because it's not easy work, it's hard work and it's very unglorious.

12:23 But at the end of the day, if we do our job right, the whole ecosystem kind of moves along smoothly and I've reduced the burden on upstream maintainers and I've gotten those fixes in the hands of people that potentially are going to be under attack by bad guys.

12:33 Yeah. And that's actually a great segue into my next question, which was about the timing of this. So, you mentioned there hasn't always been a lot of resolve or or a lot of drive to get this moving. What was it that finally pushed this over the edge? Was it a specific event or an incident that happened? If you could speak to that.

12:53 I can. We had a similar and the longer you're in technology, the more you see it's all a one big cycle. Um, and we've seen this story on a smaller scale about eight years ago when people started really getting into fuzzing and there were all these people fuzzing and throwing 700 page reports over at projects and vendors and everybody thought it was the end of the world and it wasn't. We got through it.

13:22 Fun story. We all lived and we're in a kind of a similar moment. But today instead of using scanners, people are using large language models and other AI enabled tools. And as I mentioned at the beginning of this year, the tools got pretty good at finding problems. And then like the secret advanced model stuff, the frontier model things, those actually have some interesting capabilities where they're able to make connections they aren’t good at.

13:51 They are good, but they don't only just find one problem and then move on. They're very good at finding multiple problems and then figuring out how to combine those into a bigger problem to get something like um remote code execution.

14:05 And that's where the big motivation for this is that the large language models got very good at finding problems. So, people started using these tools, whether you've got access to an advanced tool like Mythos or you've got like a 20 uh dollar $20 a month GPT account or using a local openweight model like Kimmy or Quinn or whatever, these tools are able to find these things.

14:37 They generate a ton of reports and people just start throwing them upstream. And unfortunately, every report that goes to an upstream project on average takes between two and eight hours for those developers to address. They have to understand the problem. They have to think about how it's exploited. And then they need to think about a patch. And this is all unplanned work.

15:00 You know, generally software engineers don't love maintenance. They're really thinking about features and, you know, I've got this big bug fix that's important to my business or whatever.

15:10 And they're not really thinking that it's my greatest dream to fix a bunch of low severity vulnerabilities. But that's where what caused all this and what created decrees was that there are people finding thousands of problems and they don't understand either they're bombarding upstream or they don't understand how to get these things moved upstream.

15:28 And that's where Akrites stands in. It will help make sure that those reports have all the necessary data elements to reduce that time frame for upstream maintainers so that they can maybe get a full report reviewed in maybe an hour or two and they're able to have a patch at the end of that so that they can they decide how they want to address this.

15:49 Yeah. And I think that makes a lot of sense and there's a lot of value there. Let's talk about who was all involved in this Project Akrites.

15:58 What organizations came together and are contributing their expertise to this initiative.

16:05 It's a really interesting coalition and I'm excited like within my work in the open SSF I thought we had a really wide cast of characters. Um but with the Akrites it's very interesting.

16:18 So we have the involvement of the hyperscalers who are kind of the hosts of a lot of infrastructure and they serve up a lot of this software. So they are very concerned. We have the frontier model providers, you know, Anthropic and open AI in addition to Google, Microsoft and AWS who have their own models.

16:36 We also have a large interest in the financial sector. We have several large banks that are participating. I also have some startups. So, I've got some small people that they're security researchers and they're leveraging these tools and they see value in contributing and trying to help be part of the solution.

16:50 And Sonatype is also a member here. So we're really excited.

16:56 Who are those guys?

16:58 I know those folks. Ah. Um but again it's a really interesting mix of kind of high-tech organizations, traditional brick and mortar organizations like banks and other um like telco providers.

17:13 We have our friends from Ericson who are very involved. And um at the end of the day like today we're up to 24 organizations participating and unlike other Linux Foundation where there's you know a membership fee and that goes to help supporting the organization hiring staff and putting on conferences is very different in that there is that membership fee to help pay for our infrastructure and so that we have tokens so we can run the models and whatnot but also there's a donation of technical engineers.

17:46 So we're looking for security researchers, security engineers, pentesters, product security engineers and these people are helping us review these cases. They might have relationships upstream. They might be a contributor or maintainer for a project.

18:00 So they would be kind of our gateway in but broadly they're helping us make sure that when we're ready to go upstream all those reports have a clear description. They've got a proof of concept, proof of vulnerability.

18:14 So the developer understands how to trigger the bug. Uh they have testing both positive and negative testing so that we know that here's how to find the bug. Here's how to show the bug is gone. And then most importantly we are requiring patches and I don't have the illusion that upstream is going to accept all of our patches, but at least I'm giving upstream something to start from.

18:38 They have it's they're not looking at a blank piece of paper like how do I fix this? I at least give them a solution and they might agree with it or they might write their own but at least I've kind of helped give them a jump start on their work and then we'll help when public disclosure comes around with the communications and making sure the whole ecosystem is aware and it has things like a CVE ID or it has you know severity scoring using CVSS and other methodology that type of stuff that enterprises need to manage their vulnerability programs.

19:09 Yeah. And let's dive into that a little bit more. Um, some of the some of the things that enterprises need uh when they're when they're managing their vulnerability programs.

19:24 What makes the uh Akrites project different from maybe other initiatives that are out there? And also why it matters as we look at the bigger picture. Um what is it that Akrites really brings to the table to help enterprises in their security journeys?

19:41 Well, we bring a bunch of things to the table and many of the people that are participating in the cert and part of the project we all kind of had our origin stories in some type of enterprise somewhere.

19:57 So I worked in banking, I worked in insurance and healthcare. So I have a very strong leaning towards enterprise operations.

20:03 I understand what it needs and so do many of my peers. So you know what we're bringing to the table is not only this experience and kind of that empathy for downstream but also the empathy for the developers.

20:17 I've been doing apps for 25 years or more. So I am not a developer myself but every day I work with developers. So I kind of understand how to communicate with developers and how to explain things in terms that developers can relate to.

20:31 So that's the whole team is we're very developer focused and part of our core tenants within accreties is that we respect and we are empathetic to the upstream project.

20:47 So wherever they have published rules or channels or ways that they want things reported that's the rule of the land for us. I follow what Greg KH and the colonel tells me how they want to report and that's where other groups don't necessarily have this experience.

21:00 I've been doing upstream vulnerability coordination for a long time. Uh I've helped coordinate a lot of very large ecosystem events with upstream and so we have this understanding.

21:14 Um many of the group helped write a lot of these standards. So you know I've participated in the CVE program. I've participated in the CVSs um vulnerability severity scoring system.

21:20 So we have this understanding that these are the norms. These are how a large enterprise works and I'm not trying to impose that on upstream but I'm trying to help when things go live once upstream has done their job and accepted and pushed out the patch.

21:37 I want to make sure that downstream has the things they need like a CVE ID so that they go into their vulnerability management program. It goes into static code analysis tools and other things so that you know their programs can operate and understand where in their portfolio they need to react and get these patches and a lot of people are focused on the scanning problem.

22:05 Everybody's really fixated on finding problems for upstream finding the problem isn't useful. Fixing the problem is what matters.

22:11 Yeah. And it sounds like with Project Akrites, the credibility is there because you've been in these positions. You've you've been on the development side, you've been on the appstacc side, you've been in these heavily regulated industries where they have all these compliance mandates that they need to meet with and and meet these different requirements.

22:34 Would you say that your background both you personally and the other members on this kind of project do you think that helps since you're speaking their language you've been in their shoes?

22:52 I would imagine that gives you more credibility and more uh uh a little more weight to what you have to say.

22:57 We hope so. But you know, every day we earn trust and it could be lost in a second. But yeah, you know, I and the members of the team have been involved in this upstream ecosystem.

23:12 So we understand how not all projects work the same way. There's nuances and kind of a thousand shades of gray and we understand and that's where a lot of like when you talk to government regulators, they don't open source to them is a monolith. It is a kind of contiguous thing which is absolutely false.

23:29 Every project functions differently. They have different rules, different norms. And you know again having that understanding and empathy I feel is going to help us be successful.

23:37 And you know we are beta testing right now with a core of upstream engineers.

23:44 We're also plugged in through our have by virtue of doing this type of job for so long. We know all the major players at the other foundations and the other security teams and we've reached out and we're engaging with them because I want to help enable them to go fast because at the end of the day some of us are creators but all of us are consumers of open source.

24:08 Open source is absolutely everything. And if I do my job right, I'm helping protect me and my family, you know, yay. So there's that selfishness of it, but I also help every other user of these technologies.

24:24 Everybody uses cloud. Everybody has a computer in their car, their phone, and there's just so much of the world economy that runs on open source software.

24:34 It's, you know, I feel really privileged to be in this place and have had these experiences that allow me to try to make this successfully land upstream so that we can help everybody everywhere around the world.

24:47 Yeah, and I'm glad you brought up that point on the world running on open source because I actually heard a cool stat the other day that I'm going to share here. But uh basically if open source were a country, if the usage and consumption of open source, the value behind that were the GDP of a country, it would actually make it the third richest economy in the world.

25:13 Um and that is wow. That's 8.8 trillion dollars. Um, it's the annual value of open source. So, it's wild.

25:24 It's everywhere. Like you said, it's on my watch. It's on my printer. It’s in every single thing that everyone uses every day no matter where you are.

25:34 And that's another aspect of this is that every nation on the planet is very concerned about this. there. You know, everyone's trying to think about how they're going to write laws to regulate AI, but the thing is everybody's using AI and open source is developed absolutely everywhere.

25:54 So, we kind of have to reach the developers where they live. Open source is the greatest co-engineering effort ever in existence. So, like the kernel has thousands of developers every day toiling away on this one thing. Well, this was one giant thing, but that's you again, we have to think about things in those terms.

26:13 And the scale, the magnification, if I and the team do a good job, the impact is 8 billion people that we helped, you know. Yeah, maybe we made their days a little bit better because they didn't have to worry about some bug.

26:25 Yeah, 100%. Uh well, Crobe, I I know we are running a little short on time here, but I did want to give you uh the final word here.

26:35 Uh as we're signing off, we're thinking about the future of what's coming up in the world of open source. We talked about AI and the impact that it's having in development. Um, are there any final thoughts you have for maybe best practices or advice that you would give to people who may maybe they're looking at AI, you know, and how it can help speed up development, but there's also risks that come with it.

27:03 What's a piece of advice that you would leave them with? What I would leave you with is when I talked about the three AI problems, um, the OpenSSF and CNCF, we wrote an ebook a couple months ago where we actually talked about a lot of these things and about the economics of using these large language models.

27:22 So my advice would be these tools are amazing. They can do unimaginable things, but again going back to Uncle Ben, with great power comes great responsibility.

27:33 And these things aren't going to replace you, but they can absolutely amplify your output and your ability to help. But you need to understand their limitations. They're not they're math. They're not magic.

27:47 And they need to understand how to effectively use that. And the difference between just telling a generative AI prompt, write me a website versus getting in there and using your software engineering expertise to say, I want to follow this standard and use this specification. I need to manage the amount of dependencies that are added in and kind of effectively coaching it and thinking about it as like a junior uh developer or an assistant.

28:15 You get amazing results. And you know I am not a developer but I leverage the tool like a business analyst every day where it helps me with my cyber security practice and it's again but you have to understand the limitations and also be uh skeptical of the results you know trust but verify you always have it you know show its homework.

28:35 Crobe this has been great this has been very informative for myself. I invite all of our listeners to check out the Akrites Project, which is part of Crob's last. How long have you been working on this for many years now?

28:47 I wrote the first plan five years ago. We've been actively working on this particular version of it since February. So it's been a long time coming. I'm very excited to see um accretes.dev is where you can get the public internet uh version of things. We have a public GitHub where you can check out the policies and kind of the operational things.

29:12 And eventually as we get to a point where we're happy with our tooling, all of that's going to be contributed as an open source project that people could uh fork and use for their own.

29:19 Awesome. Well, yeah, thanks again, Crob. Always appreciate uh talking with you, your level of expertise and you know, the Linux Foundation in general and and all of the good that the Linux Foundation brings to the world of open source. Uh to our viewers today, uh if this is your first time watching the open source open mic podcast, I invite you to subscribe for future episodes and leave a comment or a question. If you have a question for Crob, please leave it down below and we'll see you next time on another episode of Open Source Open Mic. Thanks.