Platform engineering usually starts with a good idea. Make things easier. Standardize the boring parts. Give developers a path that works. But eventually, every platform hits the same question. What are we actually willing to say no to? Because if every workload belongs on Kubernetes, every tool has to be centralized. And every use case has to fit the same golden path.
The platform that was supposed to accelerate teams can become the thing slowing them down. I'm Brian Teller from Teller's Tech, and this is Ship It Weekly. Welcome back to Ship It Weekly, where I filter the noise and focus on what matters when you are the one running infrastructure and owning reliability. Most weeks, it's a quick DevOps, SRE, platform, cloud, and security news recap.
In between those, I do conversation episodes with people building and operating the systems we all depend on. Today, I'm joined by Justin Garrison. Field CTO at Sidero Labs, and co-host of Fork Around and Find Out. We talk about Kubernetes moving beyond the cloud, why some companies are looking back toward bare metal and on-prem infrastructure, and where cost, sovereignty, and complexity are pushing that shift.
We also get into golden paths, why platform teams need to know what they are willing to say no to. What actually deserves to be centralized, and why Kubernetes is not automatically the right answer for every workload.
And near the end, we talk about AI, security, tooling fatigue, on-call burnout, and why chasing every new technology matters less than finding the parts of engineering you actually enjoy getting good at. All right, let's jump in. Today I'm joined by Justin Garrison. He's Field CTO at Sidero Labs, co-host of Fork Around and Find Out. And we're going to talk about Field CTO.
We're going to talk about infrastructure and whatever else comes top of mind. Justin, thank you for joining me. Yeah, thanks for having me, Brian. So explain your new title, Field CTO. Excited to hear. Yeah, it's a externally facing position at some companies. It's usually more of a tech company thing. And I'd say oftentimes it's, you know, smaller tech companies. I don't, I see it sometimes at larger ones.
But sometimes they're like segmented by region or something like that. But just like a CTO might say, this is how engineers internally should work and practices we have and resources we have. The Field CTO is kind of like an external version of that, where you're kind of helping people architect around your solutions and your software. Sometimes you're making content and you're thought leading in some cases.
But for the most part, you're just you're trying to help. People be as successful as possible with whatever it is that they're trying to use with your software. And so that's kind of a, it's a, it's a bridge between kind of sales and engineering. Sometimes there is, I do quite a bit of engineering. A lot of times it's like POCs and custom engineering for like a specific problem.
And then we see how we can fold that back into the products. But then I go and I just learn stuff from a lot of really smart people doing really hard problems and it's real fun. Awesome. Yeah, it sounds fun. I guess, what does Sidero Labs do and How are you, like, what are these POCs that you're building? Yeah, yeah. Sidero is the Greek word for iron.
And if you are familiar with infrastructure in the last 10 years, you will know that Greek words usually mean something Kubernetes and iron usually means bare metal. And so that's kind of where we fit in a lot of this is we try to focus on bare metal Kubernetes. We build an operating system called Talos Linux, which came out of a lot of frustrations with managing Kubernetes and Linux at scale.
So it stripped down as much as possible. But really, we focus on what would it take to run Kubernetes anywhere. And specifically, we like to focus on kind of the bare metal edge use cases. And where do you see, because you also worked, I guess, at Amazon on EKS. How have you seen the evolution of Kubernetes over like the last five years?
And I guess more specifically, the last year or two with everything that's happened with AI development, kind of causing everything to iterate quickly.
Yeah, I mean, Kubernetes, in many ways became productized and essentially kind of like proprietary software right and anytime you go look at what does EKS actually do what does GKE do what do these hosted services do they are all proprietary software you cannot run GKE in your data center. It is their software that you will never be able to use, right? And so it's like, oh, it's based on Kubernetes.
So it's all compatible and I can move it if I want to and portability. And it's like, eh, not really, right? Like you really have to hide all of the benefits you get out of using GKE or EKS in order to be able to do that. So that was kind of like just Kubernetes was going that route because people want the easy path. They want decisions to be made for them. They don't really want what happened with like... OpenStack.
OpenStack was every decision you had to make. It was just like, oh, well, you got to pick which version of storage do you want? Which version of networking do you want? Which version? And then it just became really hard to manage because you were piecing everything together. So proprietary stacks are great. They fit together and they're usable in certain situations. Obviously, it's still growing.
It's not slowing down. People are more and more every day. They're still going to all of those services. They're still running more Kubernetes than ever before. And then the flip side of that is people wanted to do it more places. And I'd say three or four years ago, pretty much every major vendor had some form of, you can run this on-prem, whatever they called it. I worked on EKS Anywhere.
It was kind of the last project I worked on at Amazon. And it was a, it was supposed to be a flavor of EKS that you ran in your data center. And it was nothing like what the actual backend of EKS look like. It was like, yeah, we run Kubernetes. It's like, well, what's it based on? It's like, well, we have this like fork of Kubernetes called EKS distro.
And then like, we just run that like, well, how it's like, well, it completely different, right? It's all Cluster API based. And the backend of EKS is like cube ADM and a bunch of Lambda functions. Like that's literally like, oh, you can't run that on-prem and you don't want to because. EKS does it so that they can run hundreds of thousands of them. You need 10, right?
Like you don't need the same architecture to do it that way. So same thing with like GKE had Anthos, I forget what they call it now. Azure had one, everyone had this sort of like flavor of, we want you to be able to run it yourself. And none of them were actually the same under the hood. It was always just like a Kubernetes management layer. That's most of the, almost all of them were based on Cluster API.
But that was kind of where it was heading from all, most accounts like that has failed. And almost all of them, they've all pivoted at least twice now. They've all re -architected, they've added or removed features. In many cases, they've removed headcount. So those solutions are no longer managed, as maintained as they were, right? Like, especially like Broadcom had a whole Tanzu thing and that was really cool.
And they had all these things like, and then they just fired the whole team. And they're just like, ah, we'll just keep the customers happy. And then we don't need to keep investing in that because. All of them were still venues for get people to our cloud. And we wanted to get someone to run this on-prem so that eventually they would come to us when they wanted to run it in a cloud. Makes sense.
I would say that's one of the biggest differences and one of the reasons I enjoyed Talos and Sidera was the fact that it wasn't trying to move you anywhere.
All the software was like actually you should just stay on-prem like you should just stay where you want to with the software and it does not matter we're not trying to sell you to we have a a sas version of our like management plane but like you can self -host the whole thing like we don't want you to have to move unless you want to unless you say like actually this is easier let me go ahead but like we don't have
The extra motives that a lot of those cloud providers have yeah we were talking I guess before we started as well that there's a somewhat of an insurgence again in on-prem, people moving away from cloud or moving to multi -cloud or moving cloud providers.
What do you think is the reasoning for that? There's a lot of reasons. I'd say that the two biggest ones are costs and where people are mature in their cloud environments and they're like, actually, this is really expensive.
Renting these servers it just it costs a lot of money and I I wrote the best practices that amazon for EKS the best practice for cost optimization and it basically always came down to the same thing like well you should right size your workloads and you should use reserved instances and you should try to you know like just use spot wherever you could and in those cases you know you could reduce your bill by maybe up
To like 50 60 and that's great because people like this is amazing and then some people are mature enough to get there And they've reduced their bill by 60 % and they're still like, actually, this is still more expensive than it should be when we go on -prem.
And they also have a lot of overhead for that. Autoscaling takes a lot of engineering time. It takes a lot of testing and validation. When we used to autoscale back in the day, we made a budget. We had a budget once a year. That was FinOps. We have teams of FinOps people that you're hiring to try to reduce your spend. And originally that was called a budget. And people didn't have to do this.
You didn't have to pay for extra tooling, extra people to watch it. It was just like, no, we spent the budget. And if it fit in the budget, then great. That's what we wanted. And we didn't need to spend any time on auto scaling. And all of this stuff was handled at the beginning of the year. Was it flexible? No. Was it amazing? No. But you got to ignore a lot of stuff that you can't ignore anymore. And so I see some companies going to...
On-prem or just like a simpler environment because like actually I'm auto scaling by 20 every day and it's still not saving me very much money and it's just actually causing me more pain because I have to move data whenever I auto scale and I have to like drain nodes and I have to you know all this stuff just becomes fatigue over time and so people are just like I just want like a simpler environment and then the
Other side is is honestly like political in the countries and and laws affect real people uh I'd say a lot of our customers are in Europe they're European countries or or not the united states right or just anywhere that like says like actually we're not going to use a big cloud provider from the united states because hey we we don't trust u .s companies today anymore or or we we can't because we like the size of
Just land in the world right there's different politics in different areas of land and even if if your neighboring countries are close to you and you have different rules, then you have to control the data and the services and how you use them.
And so a lot of countries that have been doing that for a very long time, like it's not something we see in the United States because we have two neighbors. And in most cases, we're fairly friendly with them. And they're also pretty far away from a lot of people in the United States. Whenever I go over to places in Europe, I'm like, wow, I've been through three countries on one train ride. And that's amazing, right? I don't get that in California.
And so just the idea of like where data lives and how it trend you know moves across wires that stuff matters a lot more and so having more control over it is important to folks and they don't want to give up the kubernetes the the nice flexibility of the system and just being able to like use modern tooling uh but they will say I'm not gonna go to amazon like amazon has like a whole like like it's a different company, right?
Like to be able to sell in the EU, like they had to spin up like these other companies. And they're like, oh, we're not Amazon, we're Amazon EU. And it's like, well, no, you're still Amazon. And so like they try to hide a lot of that stuff and say like, well, politically we're part of your country. And it's like, yeah, I get it, but I still don't trust you, right? There's still a lot of that going on.
Yeah, that's fair. So what parts of Kubernetes are essential versus what parts are self -inflicted pain? I mean, it's all a little bit self -inflicted. Anything new is going to be self -inflicted. Like you want to try when we were moving to Ansible, right? Or like config management, right? Like, did you need config management? Like, maybe, maybe not.
But in general, like throughout my career, I've seen that my own cognitive ability to like understand things breaks down around like.
Probably like 30 of the thing like it's a really low number like once I had like 30 servers it was hard for me to like go beyond that and like know what's going on right and say like oh I actually understand what these 30 servers once I'm at like 50 or 60 I'm like you know like I just need to group these into some logical grouping so that I can manage a group of them a little easier and I'd say okay well now I have
10 web servers and I have two database servers and then once I get to 30 groups then I'm like well I need to like abstract this again and in kubernetes does that really well because you can just layer on these abstractions it does get difficult to know how far down the abstractions you want to go when you're looking at processes on a server and then you say okay well I probably have thousands of these like what do I
Actually like how do I actually manage this how do I understand it but at the end of the day like people's cognitive ability is is kind of difficult to like scale beyond some limited number and so we've always been abstracting it we've always been trying to kind of how do we how do we get more I don't say value out of a person but how we let one a single person manage more things and I remember my first sysadmin job
I had a hundred linux servers and that was amazing right like that was like one person managing 100 servers and again these weren't physical services were vms because I had to logically group them into things that I could maintain and redeploy and whatnot and so but it was like that was kind of the pinnacle back in the early 2000s of like where they always thought like a you know sysadmin 100 servers per sysadmin is like the golden sort of ratio of like what you can get out of them.
And now we're looking at, okay, well now a sysadmin or platform engineer, whatever we want to call it today, doing the same sort of thing, you can probably manage, you know, maybe 10 clusters, maybe 20 Kubernetes clusters. And those will probably be in like three or four different forms of what types of workloads run on them. You're gonna have your web tier, your AI tier, your state full server tier, right?
You're gonna have these clusters that are dedicated to those things, but you're still just kind of extracting the same thing. And so in that regard, Kubernetes is great at doing that abstraction for the workloads, but it still breaks down once you say, I need to, you know, I have different versions of it, or I need, you know, I need 20 different kinds of Kubernetes clusters. Okay.
You probably need to hire someone else. Yeah. I've, I know during, in my day job, we also have the issue of everybody wanting certain teams, wanting to containerize certain applications that maybe.
Don't make sense to be containerized like long -running complicated tasks that could cause noisy neighbor issues that sort of thing that just don't make sense for kubernetes workload but it's what they're what they feel comfortable with so you know they pipe it once once you learn a tool you want to use it for everything right like people are still abusing ansible because it's just ssh in a for loop right and you're
Like actually I can just use this everywhere for everything and people were using it for application deployments was ansible ever meant for that no but It's the tool they're familiar with so they can do that.
I remember when I first was learning Go and I had some coworkers that knew a little bit of Go. And they're like, okay, well, I take my Go binary and I put it in a container. I was like, why would you do that? What is the point of putting a single, a statically compiled Go binary in a container when I can just SSH it to a server and run it?
And it took me a little while to realize that actually the benefit of these containers wasn't, it was a packaging mechanism, right? I don't care, even if it's a single binary, statically, it doesn't matter.
Just the ability to be able to, you know, push it to a registry, to replicate it a thousand times, to do all these things and to isolate it so I could run it and not have to worry about like lock files or port contention, all that stuff ends up being a really good idea, even if the thing doesn't require it.
And so like that sort of common sense around or like common knowledge on this is how this thing works and how we can use it a bunch of times ends up being pretty good for most use cases. So given that, what's your take on like golden paths? Are they useful? Like if I want to set up workloads, is there a golden path to follow? I mean, every company I've been at has their own golden path, right?
Paths end up following policies and procedures at companies. And those usually reflect some form of org chart. And so it's like, oh, depending on how your org chart looks. Depends on what your golden path is going to be. And that looks very different at a small company and a large company. And even two medium -sized companies aren't going to have the same golden path.
So it's one of the reasons I think platform engineering is so popular right now is because everyone has to have their own path. Because you can't just say, here's the path, go follow it. There are some generic tools that, hey, maybe these will help you fill in some of these gaps. But in general, you kind of have to just know.
This is what our policy is around this piece of software or when we ship things or how things get to production or the testing frameworks. And in many cases, those things also really rely on like the talent you have, right?
The people you have in charge and not just the engineering teams, but also the leadership teams being able to understand how complex something is or when you should actually send something out, right? And so it's like, if you have a really you know, forward thinking and mature engineering team.
But you have an executive who's like, no, we never ship on Friday because I've been burned too many times for shipping on Friday. Right. And that's just a thing that happens that you say, actually, our policy is we can't ship on Friday. Why? I don't know, because the exec said so. Right. It's like that. Is that part of your golden path? Sure.
Like you have a rule somewhere that there's like a, you know, you do not ship. This is Friday. You need emergency, you know, authorization from the CTO or something. Right. And like that's part of that path. And people.
Embed these things into rules and that becomes their you know how all the developers do work and those always end up being like like platforms become accelerators until they they become like balls and chains right like then they're like decelerators at some point we're just like oh we're able to move fast because everything looks the same and we can do this but then once you try to add everything into that platform you're like oh now we have too many variables and we have too many thing.
So the cognitive load of this thing is more than if we just use off the shelf parts, because all of our internal stuff isn't documented. That's how he's seen it every time. Like we've been building platforms over and over again at multiple companies and using them. And when I was at Amazon, like I, the internal tooling for Amazon, I'm sure at some point was ahead of its time.
And it took me days to get it set up to be able to do one commit into an internal repo. And I'm like, this this is garbage. Like, I wish this was just GitHub. And I understand that, like, not all those things scale to the size of something like Amazon, but also every one of those tools was proprietary and not well documented.
And I was just like, hey, this would be great if I was in an office and I sat down next to someone and they walked me through it.
But as a new remote employee trying to just understand where to even find the documentation right like that stuff all falls apart and so like those golden paths again they just they they fall apart once you get down this like too many platforms or too many options on how the platform works and any sort of like even templating right like you're like oh I'm just gonna start with like a templated file and we're just
Gonna like env you know with bash like we're just gonna like expand out the variables and at some point you look at it and you're like have logic in your templates and you're like I can't follow this anymore right and that that's around the point where you say like actually maybe we should just like fork the platform and say like this is only for this type of workload and this other one is for this type of workload
And if you can split them you'll have two platforms but they'll be simpler than a single one now that makes sense so okay we have our kubernetes cluster set up we have our golden path What, in your opinion, is the next step?
What's day two look like? What is the next most important thing when setting up an application in a cluster, having that cluster set up live, especially now with AI? Does that change? From whose perspective? Is this platform engineer perspective? Yeah. From a platform engineer, I often tell them the most important thing you can figure out is what you're going to say no to.
Figure out the things that you're like, this cluster is not for that. When I was at Disney, Disney Plus had a platform internally. And the thing we said no to was stateful workloads. And it honestly was the best decision we had ever made. Because anytime we wanted to move the cluster, it was really easy for us to do. CVEs, anything else, we're like, hey, if you need stateful stuff, you go somewhere else.
I'm sorry, this platform is not for you. And and like if you can make those decisions up front and you can say like, OK, we have what we have, you know, five monitoring systems and we know how we're going to upgrade and we know how to control like all that stuff's like just getting it set up for day one. You deploy it, it can go to production.
And at some point you're going to say, OK, what what is it that I have to say no to? And if you can figure that out for day two, you'll be in a much better place for day three, because at some point you are going to have to upgrade that cluster multiple times and you're going to have to at some again talk to.
Teams that you don't know who's running this workload and they don't care about it anymore they're off to something new right and that applies to any workload in the in the era of everyone wants to be an AI company or at least everyone wants to run AI workloads that becomes harder because AI training is very stateful and needs a lot of data and you need to figure out all of your storage your networking performance
Tuning And as soon as you get all that, you're like, actually, this is going to run for a week at a time to do some sort of like fine tuning or training or something like that.
And like now I need to figure out how to snapshot it. Right. And I was like, OK. And these are these aren't new problems. These are problems that HPC environments have had for a very long time. And so a lot of what Kubernetes was built for, built for stateless web applications originally, is now running head head into this like job management and long running jobs that it just it.
Isn't well suited for at like the lower levels and and that's why people still use a lot of Slurm and they still use a lot of custom tooling for hpc and like strictly hpc environments and in kubernetes it's getting there at some point but I don't know I don't see that It's not a perfect fit.
You're going to have to layer on some other abstraction to get the features you want because Kubernetes at the API level is like you can plug in whatever you want and then go build it. And they're kind of done putting in like core functionality into the API. So given Kubernetes adoption, given saying no to things, what's a practice outside of that you wish more teams would have adopted before they needed it?
As far as like the structure of their, from a platform engineering perspective, as far as structuring for their team. It's something that I don't think the platform teams can't give themselves, but I wish that the organization would give the platform team, the agency to have like some financial power to say, to control what's going on. Right.
So if I look at like SRE, SRE, when it came out, the Google book and all this stuff, they're like, Oh, this is, this is how you do SRE. This is, this is what S and everyone's like, Oh, well it's like engineering and production and uptime and all this stuff.
It's all true, but also the thing I took out from that book was the thing that SRE at Google did that no other SRE team ever did that I saw was the SRE teams had the ability to say no. The SRE team could say, your application sucks. It has too much technical debt. I'm not going to maintain it. And they could walk away.
And then the team was on their own to like bring it up to a level of, yeah, we have to be able to fix this now. And that was only because the SRE organization had that ability, right? The funding for SRE was taken out of engineering budget from the application team. So if you wanted an SRE on your team, you had to pay for it. And then the SRE team could say, no, that's not worth our time. Right.
We have other places that we should be spending our time. Most platforms are not funded that way. Most platform teams don't get their funding from the engineering team of the application. Right. They're getting their funding from infrastructure budgets or they're getting their funding from maybe security teams. And those are the places that are saying, hey, I'm going to pay for this team.
This is a 10 person team for building our platform. Why are you building the platform? Because we need a platform. Tell me what you're trying to do. We're trying to accelerate developers so they can have less cognitive load about, developers shouldn't learn Kubernetes, we should just subtract. I'm like, well, what are you actually trying to do here? And they're like, well, we want to centralize things.
Okay, what is the benefit of you centralizing this golden path? And from my experience, the main things that you want to centralize are security, like being able to audit your systems and say like, hey, is this going to be a problem for us? If there's a CVE, you need to be able to patch it quickly. If there's a leak or a hack or something, you need to be able to audit your systems and understand what's going on.
So things like logging, monitoring, things like just like SBOMs and like security scanning, that's something you need to centralize. And if you're in a cloud environment, you need to centralize your costs, right? Because like, that's just where are we spending more money than we should be? That's like the basics of a platform is usually those.
Two things to me if you don't need to build up the golden path for all this engineering stuff if you just give them hey everyone now gets automatic data dog logging here you go and everyone automatically is in these AWS accounts and we're going to take care of the how the costs bubble up into our orgs so we can see that at central place and everyone's going to get you have to have this s -bomb in place in order to
Ship to production right that's the basics of like a platform team and I wish more platform teams or more organizations would understand you don't want to centralize everything.
There's actually not a lot of value in putting everything in the same centralized tool, because again, that's just going to slow you down at some point. Most teams are going to say, we got to centralize CI/CD, because back in the day, Jenkins was the way to go, and it was really hard to set up Jenkins servers, and so you had to set up one.
And that one became so critical to how everything worked in your company that no one could upgrade it. And so then you had this really old Jenkins server until the next team said, you know what? I'm going to send up my own Jenkins server because the other one sucks. And then you had this like battle of Jenkins servers in most companies. And at some point, everyone's like, I like the new Jenkins server.
It has the new UI or it does the new, you know, whatever, right? Like it's we want to move to the new CI/CD. It's like, well, like. Yeah, it's a cost to maintain those sorts of things. And it is valuable to provide an optional, hey, you can use this template that we have that allows you to do something faster. But mandating centralization through some sort of platform team ends up being a hindrance.
And most organizations that I talk to and platform teams that I talk to, they don't understand that. They just say, we just have to centralize everything. We need everything to run on Kubernetes. I'm like, no, you don't. Trust me, you don't want everything.
And also like, this this keep this notion of like you need kubernetes clusters to do anything is that was like before AI right now we need AI to do anything like before anyone's gonna write any code like well if claude's down I can't write any code like what are you doing like do you do you remember how this works like you you can't is it worth your time I don't know you have to decide that right like how slow are
You going to be without it should you take a break for an hour do you need to go refill your tokens whatever right but like Back in the day, like we keep doing the same thing just with a different technology, right?
Because even like 10, 15 years ago, you put a credit card on a cloud provider and that accelerated what you could do because it took too long to get a server in the data center. And then later you're like, okay, well now I'm going to, you know, my credit card is going to go to something else, some other service. So I don't have to run this thing anymore because I just don't have the time to.
And now credit cards are on cloud. And that's how like the amount of people I know that are paying like the $200, you know, thing, just so they're like, I just need to stay like up to not up to date. And even it's just like, I need to be. As productive as the next person, I'm like, no, like, what are you doing? Like, let's just understand where the problem actually lies. And yeah, we're going to keep doing this.
There's going to be a new technology that's going to come out in five years. People are going to put their credit card on it and they're going to say, actually, I just, I have to have this to work. And I always sort of like reevaluate that. Like, what are you actually trying to get done? And when does it actually need to be done by? And then once we have that conversation, we can say, okay.
Well, then let's plan this, right? It's just like a budget again. If I know my workload is going to scale to maximum of this in the next 12 months, I can plan for that. And people are so against planning today that they're just like, I just need to spend money to solve the problems. So do you think that AI kind of causes people to do less critical thinking on their own?
And I mean, the Dora report last year says that we see this AI productivity gain. Across the industry, but we're not able to quantify it correctly or quantify it accurately in the data analytics that we get out of our actual number of commits or number of features that are released or whatever.
So do you think that there is this inherent symbiotic relationship that people have formed now with their AI as far as engineers unable to be engineers without their AI tooling? I think engineers have formed too close of... Relationships with a lot of their tools. It's not an AI problem.
Vim and Emacs, this is a problem that people have religious debates about for a very long time about pretty much any tooling because they feel attached and they identify with those things. And AI is another version of that.
I think the fact that AI is changing so frequently that people are having a hard time even even figuring out what their identity is because people that were co -pilot advocates last year right if you were out there today saying like co -pilot's the way to go everyone's gonna laugh at you like no way man why are you you gotta be on clod like what are you doing right and this was this was literally eight months ago
That people were just like this was the way to go and in microsoft just like everything's gonna be co -pilot and now it's like they kind of lost the open AI you know deals and now they're just like well now we're just gonna run whatever AI llm we think is the right one and and so like Everyone's changing their mind about this stuff.
I do think that the kind of the I guess the status quo of how people are working today where I want to review code more than write code is more common in some of these tools. I don't think that's and I do think AI has accelerated that but I don't think it's like completely new. Right. Like I always think of Ruby on Rails was amazing. The first time I ran it, where I was just like, oh yeah, make me a new website.
And like, it would template out like 20 folders and all this stuff. And it's like, you just put in these eight variables and we're ready to go. I'm like, that's amazing. Right. But like, and I didn't review the code. I didn't look at it. I just said, okay, well, yeah, Ruby on Rails, ship it. Like, you know what you're doing. And we're, we're kind of, it's, it's that, but it's.
A lot more flexible right like we're just we're doing the same sort of stuff and even like go back a long time right compilers and compilers they just take your you know take your high level cobalt and and make it machine code and that's amazing right like you can actually do that without needing to review it and at some point people move more and more to like how do this architect together where where are the
Important things for me to have more input versus just being able to know like oh I know how these pieces should fit together someone else can actually like make the glue or the templates for me So do you think that with the advent of what we have, Claude, Mythos, and Project Glasswing, do you think that that changes the prioritization now more towards cybersecurity on four platform teams and DevOps teams alike?
Or do you think that it's status quo and we just work from the outside in, make sure there's a WAF, lock everything down, and make sure our packages are up to date? And that's enough. I love how optimistic you are that people are actually using the WAFs and everything. I mean, because I look at what people are shipping today.
None of that is in place right this is just like oh yeah we're just gonna we're just gonna burn through this as fast as we can and again and that's I mean to some of the larger you know companies that have more to lose they have the policies they have the practices they're gonna say like actually no you can't ship that to production without these things and engineers and people are pretty smart they're gonna try to
Work around it and figure out how can I get my thing done to be successful my way everyone always tries to work around some of that stuff but at the end of the day like there is going to be some controls on okay Where is this actually going?
What does it have access to? I think all those attacks are going to be more sophisticated, but I don't see it being measured yet. Like, S -bombs are measured, right? Security is always reactionary in that, like, oh, there's no CVEs in my architecture. I'm like, yeah, there are. It's like, no one knows about them yet, right? And I think that that is going to be, until we can measure it, we can't manage it.
And so we have to flip this to be a little more proactive. And WAFs are one way of doing that. We're just going to have a firewall, we're going to say no DDoS, whatever. We're going to put some things in place, but once your application's online, it's fair game for anyone to start poking at it. And that is still, I see, who was it just last week? Calendly? They closed source their software.
Actually, we don't want people to see our code. We now need to close source it because being open source is a vector for... An AI bot to read our code and say, oh, I see exactly where this is. At the end of the day, it's probably going to be some subscription and it's going to be tied to some level of insurance, right?
There's cybersecurity insurance and some of those cybersecurity insurances are going to say, hey, you have to run this mythos on your code base once a month or something like that. There's going to be some requirements there just to say, we did what we could. But at the end of the day, you still have to... If you don't have a way for customers to access your stuff, it's going to be really hard to sell something.
No, that's for sure. Okay. So closing out, what's one piece of advice for platform or SRE people trying to stay sane this year with everything that's happening, all of the new features, all of the new models, how do we stay sane? I would say find the things that you're interested in, right? Like don't just keep throwing YAML at the wall, right? Like, like you actually like, Figure out, oh, you know what?
I really like networking. I think storage is cool. I want, you know, I want to understand some level of something like find the things that you're passionate about and that you think, oh, that's interesting. Whether it's a career path or not, right? Like whether it's like, you don't have to plan out. I think this is going to be big, so I'm going to go into it, right? Let's just like figure out for you.
Like when I first started Kubernetes, it was not the winner, right? Like this was, this was back in like Mezos was the thing that everyone ran. And when I was like, oh.
This kubernetes thing is kind of cool and why is it different than what Mesos offers or what Slurm offers or and and just knowing that like it doesn't always work out I have tons of projects that were just throw away that like I learned a skill deep enough that I don't need it anymore but it doesn't matter that it didn't pay out for me like I I had fun I learned the thing and in having joy in what you do and is more important than like having some form of like financial success, right?
Like the happiness, the joy, the being able to like still be yourself is a lot harder. And platform teams have a lot of breadth of things that they're going to look at. SRE teams, you're going to burn out no matter what. Like there's, oh, if you are on call, I don't care where you're on call for, that level of pressure is always going to cause some level of burnout.
So talk to your manager early about, hey, how can I make sure that like this is sustainable at a team level? Like, hey, we know we're going to lose one person a year to this. How do we handle on-call rotations when we're short people? Those are all things that you should be thinking about on making this a little more enjoyable for yourself and not just burning all your tokens and then going home and just crashing.
At some point, you're just like, what do you have fun doing? And what do you look forward to every week? If there's something at work that you look forward to, you should do more of that. Right. You should you should say, like, actually, you know, I really like this thing. When can I do that twice a week now? Can I do it? Maybe, you know, like like dedicate a week, a month or something.
That's the important stuff, because that's the stuff that will give you the energy just to be a better.
Person in general like you are more enjoyable to be around when you like what you're doing in a lot of cases and you like doing it with the same you know right people and that's not always a choice everyone can make it's not always a freedom that everyone has but when you do have those things try to do more of it and try to focus on it I 100 agree with that justin where can people find your work I'm on social media
Quite a bit I'm hanging out a lot on blue sky um justin garrison.com is my website I still blog there I've been blogging for over 20 years and over that the entire 20 years I've I've averaged one blog post a month awesome which is amazing right like I look back at it when I migrated from a bunch of different places into you know into my site and I was like wow like I have a lot of articles like that's pretty crazy
And I just added them up and I'm like actually that's that's amazing like I don't blog all the time but I when I do I go into like I'm gonna do three or four blog posts and I'll put them all out there.
And so I'm on average, I'm about one a month. And then I also, in the podcast Workaround Findouts, we're doing that once a month as well. So I kind of have those two things that I just try to keep up. And those are where I'm interested in things and when I want to talk to people and where I find joy. Awesome. Well, I'll put links for all of that in the show notes. And Justin, thank you so much for coming on.
Really appreciate it. Yeah. Thank you. That was my conversation with Justin Garrison, Field CTO at Sidero Labs. And co-host of Fork Around and Find Out. The biggest takeaway for me is that good platforms are defined as much by what they refuse to do as what they support. It is tempting to centralize everything. One Kubernetes platform. One CI/CD system. One golden path. One way every team is supposed to work.
And that can absolutely make things easier at first. But eventually... The exceptions pile up. The platform absorbs more responsibility. And abstraction that was supposed to reduce cognitive load becomes its own source of cognitive load. I liked Justin's example from Disney +, where the platform simply said no to stateful workloads. That constraint made everything else easier to operate, upgrade, move, and recover.
That is a useful way to think about platform engineering in general. Do not start with how do we support everything? Start with what problem are we actually solving? What should be centralized? And what should explicitly live somewhere else? The same idea applies to Kubernetes and AI. Once we learn a useful tool, we tend to want to use it everywhere. Kubernetes becomes the answer to every workload.
AI becomes a requirement before somebody writes a line of code. Sometimes that is useful. Sometimes we are just replacing one dependency with another. So my takeaway is to keep questioning the abstraction. Know why you are using Kubernetes. Know why you are centralizing something. Know what your golden path is optimizing for. And know when the right answer is simply this platform is not for that.
You can find Justin's work, his writing, and fork around and find out through the links in the show notes. Follow or subscribe to Ship It Weekly wherever you listen and find previous episodes at shipitweekly.fm. Thanks for listening, and I'll see you later this week.