0:00
Platform engineering usually starts with a good
0:03
idea. Make things easier. Standardize the boring
0:07
parts. Give developers a path that works. But
0:12
eventually, every platform hits the same question.
0:16
What are we actually willing to say no to? Because
0:20
if every workload belongs on Kubernetes, every
0:23
tool has to be centralized. And every use case
0:27
has to fit the same golden path. The platform
0:30
that was supposed to accelerate teams can become
0:34
the thing slowing them down. I'm Brian Teller
0:37
from Teller's Tech, and this is Ship It Weekly.
0:57
Welcome back to Ship It Weekly, where I filter
1:00
the noise and focus on what matters when you
1:03
are the one running infrastructure and owning
1:06
reliability. Most weeks, it's a quick DevOps,
1:09
SRE, platform, cloud, and security news recap.
1:14
In between those, I do conversation episodes
1:17
with people building and operating the systems
1:21
we all depend on. Today, I'm joined by Justin
1:25
Garrison. Field CTO at Sidero Labs, and co-host
1:30
of Fork Around and Find Out. We talk about Kubernetes
1:34
moving beyond the cloud, why some companies are
1:37
looking back toward bare metal and on-prem infrastructure,
1:41
and where cost, sovereignty, and complexity are
1:46
pushing that shift. We also get into golden paths,
1:50
why platform teams need to know what they are
1:53
willing to say no to. what actually deserves
1:56
to be centralized, and why Kubernetes is not
2:00
automatically the right answer for every workload.
2:04
And near the end, we talk about AI, security,
2:08
tooling fatigue, on-call burnout, and why chasing
2:13
every new technology matters less than finding
2:16
the parts of engineering you actually enjoy getting
2:20
good at. All right, let's jump in. Today I'm
2:28
joined by Justin Garrison. He's Field CTO at
2:30
Sidero Labs, co-host of Fork Around and Find
2:33
Out. And we're going to talk about Field CTO.
2:35
We're going to talk about infrastructure and
2:37
whatever else comes top of mind. Justin, thank
2:40
you for joining me. Yeah, thanks for having me,
2:41
Brian. So explain your new title, Field CTO.
2:45
Excited to hear. Yeah, it's a externally facing
2:47
position at some companies. It's usually more
2:50
of a tech company thing. And I'd say oftentimes
2:52
it's, you know, smaller tech companies. I don't,
2:54
I see it sometimes at larger ones. But sometimes
2:57
they're like segmented by region or something
3:00
like that. But just like a CTO might say, this
3:02
is how engineers internally should work and practices
3:05
we have and resources we have. The Field CTO
3:08
is kind of like an external version of that,
3:10
where you're kind of helping people architect
3:12
around your solutions and your software. Sometimes
3:15
you're making content and you're thought leading
3:19
in some cases. But for the most part, you're
3:22
just you're trying to help. people be as successful
3:25
as possible with whatever it is that they're
3:27
trying to use with your software. And so that's
3:29
kind of a, it's a, it's a bridge between kind
3:31
of sales and engineering. Sometimes there is,
3:33
I do quite a bit of engineering. A lot of times
3:35
it's like POCs and custom engineering for like
3:38
a specific problem. And then we see how we can
3:41
fold that back into the products. But then I
3:43
go and I just learn stuff from a lot of really
3:45
smart people doing really hard problems and it's
3:47
real fun. Awesome. Yeah, it sounds fun. I guess,
3:49
what does Sidero Labs do and How are you, like,
3:53
what are these POCs that you're building? Yeah,
3:55
yeah. Sidero is the Greek word for iron. And
3:58
if you are familiar with infrastructure in the
4:00
last 10 years, you will know that Greek words
4:02
usually mean something Kubernetes and iron usually
4:05
means bare metal. And so that's kind of where
4:07
we fit in a lot of this is we try to focus on
4:10
bare metal Kubernetes. We build an operating
4:12
system called Talos Linux, which came out of
4:15
a lot of frustrations with managing Kubernetes
4:17
and Linux at scale. So it stripped down as much
4:21
as possible. But really, we focus on what would
4:24
it take to run Kubernetes anywhere. And specifically,
4:26
we like to focus on kind of the bare metal edge
4:29
use cases. And where do you see, because you
4:33
also worked, I guess, at Amazon on EKS. How have
4:37
you seen the evolution of Kubernetes over like
4:39
the last five years? And I guess more specifically,
4:42
the last year or two with everything that's happened
4:45
with AI development, kind of causing everything
4:47
to iterate quickly. Yeah, I mean, Kubernetes,
4:51
in many ways became productized and essentially
4:56
kind of like proprietary software right and anytime
4:59
you go look at what does EKS actually do what
5:01
does GKE do what do these hosted services do
5:04
they are all proprietary software you cannot
5:06
run GKE in your data center. It is their software
5:09
that you will never be able to use, right? And
5:11
so it's like, oh, it's based on Kubernetes. So
5:12
it's all compatible and I can move it if I want
5:15
to and portability. And it's like, eh, not really,
5:17
right? Like you really have to hide all of the
5:20
benefits you get out of using GKE or EKS in order
5:23
to be able to do that. So that was kind of like
5:25
just Kubernetes was going that route because
5:27
people want the easy path. They want decisions
5:29
to be made for them. They don't really want what
5:31
happened with like... OpenStack. OpenStack was
5:34
every decision you had to make. It was just like,
5:36
oh, well, you got to pick which version of storage
5:39
do you want? Which version of networking do you
5:40
want? Which version? And then it just became
5:42
really hard to manage because you were piecing
5:44
everything together. So proprietary stacks are
5:47
great. They fit together and they're usable in
5:50
certain situations. Obviously, it's still growing.
5:53
It's not slowing down. People are more and more
5:56
every day. They're still going to all of those
5:58
services. They're still running more Kubernetes
5:59
than ever before. And then the flip side of that
6:02
is people wanted to do it more places. And I'd
6:05
say three or four years ago, pretty much every
6:07
major vendor had some form of, you can run this
6:10
on-prem, whatever they called it. I worked on
6:12
EKS Anywhere. It was kind of the last project
6:14
I worked on at Amazon. And it was a, it was supposed
6:17
to be a flavor of EKS that you ran in your data
6:19
center. And it was nothing like what the actual
6:21
backend of EKS look like. It was like, yeah,
6:24
we run Kubernetes. It's like, well, what's it
6:26
based on? It's like, well, we have this like
6:27
fork of Kubernetes called EKS distro. And then
6:30
like, we just run that like, well, how it's like,
6:31
well, it completely different, right? It's all
6:33
Cluster API based. And the backend of EKS is
6:36
like cube ADM and a bunch of Lambda functions.
6:38
Like that's literally like, oh, you can't run
6:40
that on-prem and you don't want to because.
6:43
EKS does it so that they can run hundreds of
6:45
thousands of them. You need 10, right? Like you
6:49
don't need the same architecture to do it that
6:51
way. So same thing with like GKE had Anthos,
6:55
I forget what they call it now. Azure had one,
6:58
everyone had this sort of like flavor of, we
7:00
want you to be able to run it yourself. And none
7:02
of them were actually the same under the hood.
7:04
It was always just like a Kubernetes management
7:06
layer. That's most of the, almost all of them
7:08
were based on Cluster API. But that was kind
7:10
of where it was heading from all, most accounts
7:13
like that has failed. And almost all of them,
7:15
they've all pivoted at least twice now. They've
7:17
all re -architected, they've added or removed
7:19
features. In many cases, they've removed headcount.
7:22
So those solutions are no longer managed, as
7:24
maintained as they were, right? Like, especially
7:26
like Broadcom had a whole Tanzu thing and that
7:29
was really cool. And they had all these things
7:31
like, and then they just fired the whole team.
7:32
And they're just like, ah, we'll just keep the
7:33
customers happy. And then we don't need to keep
7:35
investing in that because. All of them were still
7:38
venues for get people to our cloud. And we wanted
7:41
to get someone to run this on-prem so that eventually
7:44
they would come to us when they wanted to run
7:47
it in a cloud. Makes sense. I would say that's
7:50
one of the biggest differences and one of the
7:53
reasons I enjoyed Talos and Sidera was the fact
7:55
that it wasn't trying to move you anywhere. all
7:59
the software was like actually you should just
8:01
stay on-prem like you should just stay where
8:03
you want to with the software and it does not
8:05
matter we're not trying to sell you to we have
8:08
a a sas version of our like management plane
8:11
but like you can self -host the whole thing like
8:12
we don't want you to have to move unless you
8:14
want to unless you say like actually this is
8:16
easier let me go ahead but like we don't have
8:18
the extra motives that a lot of those cloud providers
8:20
have yeah we were talking i guess before we started
8:24
as well that there's a somewhat of an insurgence
8:28
again in on-prem, people moving away from cloud
8:32
or moving to multi -cloud or moving cloud providers.
8:35
What do you think is the reasoning for that?
8:38
There's a lot of reasons. I'd say that the two
8:40
biggest ones are costs and where people are mature
8:44
in their cloud environments and they're like,
8:46
actually, this is really expensive. renting these
8:49
servers it just it costs a lot of money and i
8:51
i wrote the best practices that amazon for EKS
8:54
the best practice for cost optimization and it
8:57
basically always came down to the same thing
8:58
like well you should right size your workloads
8:59
and you should use reserved instances and you
9:02
should try to you know like just use spot wherever
9:04
you could and in those cases you know you could
9:07
reduce your bill by maybe up to like 50 60 and
9:11
that's great because people like this is amazing
9:12
and then some people are mature enough to get
9:15
there And they've reduced their bill by 60 %
9:17
and they're still like, actually, this is still
9:18
more expensive than it should be when we go on
9:21
-prem. And they also have a lot of overhead for
9:24
that. Autoscaling takes a lot of engineering
9:27
time. It takes a lot of testing and validation.
9:31
When we used to autoscale back in the day, we
9:34
made a budget. We had a budget once a year. That
9:38
was FinOps. We have teams of FinOps people that
9:41
you're hiring to try to reduce your spend. And
9:44
originally that was called a budget. And people
9:46
didn't have to do this. You didn't have to pay
9:48
for extra tooling, extra people to watch it.
9:50
It was just like, no, we spent the budget. And
9:51
if it fit in the budget, then great. That's what
9:53
we wanted. And we didn't need to spend any time
9:55
on auto scaling. And all of this stuff was handled
9:58
at the beginning of the year. Was it flexible?
9:59
No. Was it amazing? No. But you got to ignore
10:02
a lot of stuff that you can't ignore anymore.
10:05
And so I see some companies going to... on-prem
10:08
or just like a simpler environment because like
10:10
actually i'm auto scaling by 20 every day and
10:14
it's still not saving me very much money and
10:16
it's just actually causing me more pain because
10:18
i have to move data whenever i auto scale and
10:21
i have to like drain nodes and i have to you
10:23
know all this stuff just becomes fatigue over
10:25
time and so people are just like i just want
10:26
like a simpler environment and then the other
10:28
side is is honestly like political in the countries
10:33
and and laws affect real people uh i'd say a
10:38
lot of our customers are in Europe they're European
10:42
countries or or not the united states right or
10:45
just anywhere that like says like actually we're
10:46
not going to use a big cloud provider from the
10:48
united states because hey we we don't trust u
10:51
.s companies today anymore or or we we can't
10:54
because we like the size of just land in the
10:58
world right there's different politics in different
11:01
areas of land and even if if your neighboring
11:03
countries are close to you and you have different
11:07
rules, then you have to control the data and
11:10
the services and how you use them. And so a lot
11:12
of countries that have been doing that for a
11:14
very long time, like it's not something we see
11:16
in the United States because we have two neighbors.
11:19
And in most cases, we're fairly friendly with
11:22
them. And they're also pretty far away from a
11:24
lot of people in the United States. Whenever
11:26
I go over to places in Europe, I'm like, wow,
11:28
I've been through three countries on one train
11:30
ride. And that's amazing, right? I don't get
11:32
that in California. And so just the idea of like
11:36
where data lives and how it trend you know moves
11:39
across wires that stuff matters a lot more and
11:42
so having more control over it is important to
11:44
folks and they don't want to give up the kubernetes
11:47
the the nice flexibility of the system and just
11:50
being able to like use modern tooling uh but
11:53
they will say i'm not gonna go to amazon like
11:56
amazon has like a whole like like it's a different
12:00
company, right? Like to be able to sell in the
12:02
EU, like they had to spin up like these other
12:04
companies. And they're like, oh, we're not Amazon,
12:05
we're Amazon EU. And it's like, well, no, you're
12:07
still Amazon. And so like they try to hide a
12:10
lot of that stuff and say like, well, politically
12:11
we're part of your country. And it's like, yeah,
12:14
I get it, but I still don't trust you, right?
12:16
There's still a lot of that going on. Yeah, that's
12:19
fair. So what parts of Kubernetes are essential
12:24
versus what parts are self -inflicted pain? I
12:29
mean, it's all a little bit self -inflicted.
12:31
Anything new is going to be self -inflicted.
12:35
Like you want to try when we were moving to Ansible,
12:38
right? Or like config management, right? Like,
12:40
did you need config management? Like, maybe,
12:42
maybe not. But in general, like throughout my
12:47
career, I've seen that my own cognitive ability
12:49
to like understand things breaks down around
12:52
like. probably like 30 of the thing like it's
12:55
a really low number like once i had like 30 servers
12:57
it was hard for me to like go beyond that and
13:00
like know what's going on right and say like
13:03
oh i actually understand what these 30 servers
13:05
once i'm at like 50 or 60 i'm like you know like
13:08
i just need to group these into some logical
13:10
grouping so that i can manage a group of them
13:13
a little easier and i'd say okay well now i have
13:14
10 web servers and i have two database servers
13:17
and then once i get to 30 groups then i'm like
13:20
well i need to like abstract this again and in
13:22
kubernetes does that really well because you
13:25
can just layer on these abstractions it does
13:27
get difficult to know how far down the abstractions
13:30
you want to go when you're looking at processes
13:32
on a server and then you say okay well i probably
13:35
have thousands of these like what do i actually
13:37
like how do i actually manage this how do i understand
13:39
it but at the end of the day like people's cognitive
13:41
ability is is kind of difficult to like scale
13:43
beyond some limited number and so we've always
13:47
been abstracting it we've always been trying
13:48
to kind of how do we how do we get more i don't
13:51
say value out of a person but how we let one
13:53
a single person manage more things and i remember
13:56
my first sysadmin job i had a hundred linux servers
13:58
and that was amazing right like that was like
14:01
one person managing 100 servers and again these
14:05
weren't physical services were vms because i
14:07
had to logically group them into things that
14:09
i could maintain and redeploy and whatnot and
14:11
so but it was like that was kind of the pinnacle
14:13
back in the early 2000s of like where they always
14:16
thought like a you know sysadmin 100 servers
14:19
per sysadmin is like the golden sort of ratio
14:21
of like what you can get out of them. And now
14:23
we're looking at, okay, well now a sysadmin or
14:25
platform engineer, whatever we want to call it
14:27
today, doing the same sort of thing, you can
14:29
probably manage, you know, maybe 10 clusters,
14:32
maybe 20 Kubernetes clusters. And those will
14:35
probably be in like three or four different forms
14:37
of what types of workloads run on them. You're
14:39
gonna have your web tier, your AI tier, your
14:41
state full server tier, right? You're gonna have
14:43
these clusters that are dedicated to those things,
14:45
but you're still just kind of extracting the
14:47
same thing. And so in that regard, Kubernetes
14:50
is great at doing that abstraction for the workloads,
14:54
but it still breaks down once you say, I need
14:56
to, you know, I have different versions of it,
14:59
or I need, you know, I need 20 different kinds
15:01
of Kubernetes clusters. Okay. You probably need
15:03
to hire someone else. Yeah. I've, I know during,
15:07
in my day job, we also have the issue of everybody
15:11
wanting certain teams, wanting to containerize
15:14
certain applications that maybe. don't make sense
15:17
to be containerized like long -running complicated
15:20
tasks that could cause noisy neighbor issues
15:22
that sort of thing that just don't make sense
15:24
for kubernetes workload but it's what they're
15:26
what they feel comfortable with so you know they
15:29
pipe it once once you learn a tool you want to
15:31
use it for everything right like people are still
15:33
abusing ansible because it's just ssh in a for
15:37
loop right and you're like actually i can just
15:38
use this everywhere for everything and people
15:39
were using it for application deployments was
15:42
ansible ever meant for that no but It's the tool
15:45
they're familiar with so they can do that. I
15:47
remember when I first was learning Go and I had
15:49
some coworkers that knew a little bit of Go.
15:51
And they're like, okay, well, I take my Go binary
15:53
and I put it in a container. I was like, why
15:55
would you do that? What is the point of putting
15:57
a single, a statically compiled Go binary in
16:00
a container when I can just SSH it to a server
16:03
and run it? And it took me a little while to
16:05
realize that actually the benefit of these containers
16:07
wasn't, it was a packaging mechanism, right?
16:10
I don't care, even if it's a single binary, statically,
16:13
it doesn't matter. Just the ability to be able
16:15
to, you know, push it to a registry, to replicate
16:17
it a thousand times, to do all these things and
16:19
to isolate it so I could run it and not have
16:21
to worry about like lock files or port contention,
16:23
all that stuff ends up being a really good idea,
16:27
even if the thing doesn't require it. And so
16:30
like that sort of common sense around or like
16:33
common knowledge on this is how this thing works
16:36
and how we can use it a bunch of times ends up
16:39
being pretty good for most use cases. So given
16:44
that, what's your take on like golden paths?
16:47
Are they useful? Like if I want to set up workloads,
16:52
is there a golden path to follow? I mean, every
16:57
company I've been at has their own golden path,
16:59
right? Paths end up following policies and procedures
17:03
at companies. And those usually reflect some
17:07
form of org chart. And so it's like, oh, depending
17:10
on how your org chart looks. Depends on what
17:12
your golden path is going to be. And that looks
17:14
very different at a small company and a large
17:17
company. And even two medium -sized companies
17:20
aren't going to have the same golden path. So
17:22
it's one of the reasons I think platform engineering
17:24
is so popular right now is because everyone has
17:26
to have their own path. Because you can't just
17:28
say, here's the path, go follow it. There are
17:30
some generic tools that, hey, maybe these will
17:33
help you fill in some of these gaps. But in general,
17:36
you kind of have to just know. this is what our
17:38
policy is around this piece of software or when
17:42
we ship things or how things get to production
17:44
or the testing frameworks. And in many cases,
17:47
those things also really rely on like the talent
17:50
you have, right? The people you have in charge
17:52
and not just the engineering teams, but also
17:54
the leadership teams being able to understand
17:56
how complex something is or when you should actually
18:00
send something out, right? And so it's like,
18:02
if you have a really you know, forward thinking
18:05
and mature engineering team. But you have an
18:08
executive who's like, no, we never ship on Friday
18:10
because I've been burned too many times for shipping
18:12
on Friday. Right. And that's just a thing that
18:14
happens that you say, actually, our policy is
18:16
we can't ship on Friday. Why? I don't know, because
18:18
the exec said so. Right. It's like that. Is that
18:20
part of your golden path? Sure. Like you have
18:23
a rule somewhere that there's like a, you know,
18:25
you do not ship. This is Friday. You need emergency,
18:27
you know, authorization from the CTO or something.
18:30
Right. And like that's part of that path. And
18:32
people. embed these things into rules and that
18:36
becomes their you know how all the developers
18:38
do work and those always end up being like like
18:43
platforms become accelerators until they they
18:45
become like balls and chains right like then
18:48
they're like decelerators at some point we're
18:49
just like oh we're able to move fast because
18:51
everything looks the same and we can do this
18:53
but then once you try to add everything into
18:54
that platform you're like oh now we have too
18:56
many variables and we have too many thing. So
18:59
the cognitive load of this thing is more than
19:02
if we just use off the shelf parts, because all
19:04
of our internal stuff isn't documented. That's
19:07
how he's seen it every time. Like we've been
19:08
building platforms over and over again at multiple
19:10
companies and using them. And when I was at Amazon,
19:13
like I, the internal tooling for Amazon, I'm
19:16
sure at some point was ahead of its time. And
19:18
it took me days to get it set up to be able to
19:20
do one commit into an internal repo. And I'm
19:22
like, this this is garbage. Like, I wish this
19:24
was just GitHub. And I understand that, like,
19:27
not all those things scale to the size of something
19:29
like Amazon, but also every one of those tools
19:31
was proprietary and not well documented. And
19:35
I was just like, hey, this would be great if
19:36
I was in an office and I sat down next to someone
19:38
and they walked me through it. but as a new remote
19:41
employee trying to just understand where to even
19:43
find the documentation right like that stuff
19:45
all falls apart and so like those golden paths
19:47
again they just they they fall apart once you
19:50
get down this like too many platforms or too
19:54
many options on how the platform works and any
19:56
sort of like even templating right like you're
19:58
like oh i'm just gonna start with like a templated
20:00
file and we're just gonna like env you know with
20:04
bash like we're just gonna like expand out the
20:06
variables and at some point you look at it and
20:08
you're like have logic in your templates and
20:10
you're like i can't follow this anymore right
20:11
and that that's around the point where you say
20:13
like actually maybe we should just like fork
20:15
the platform and say like this is only for this
20:17
type of workload and this other one is for this
20:19
type of workload and if you can split them you'll
20:22
have two platforms but they'll be simpler than
20:25
a single one now that makes sense so okay we
20:29
have our kubernetes cluster set up we have our
20:31
golden path What, in your opinion, is the next
20:34
step? What's day two look like? What is the next
20:38
most important thing when setting up an application
20:42
in a cluster, having that cluster set up live,
20:45
especially now with AI? Does that change? From
20:48
whose perspective? Is this platform engineer
20:49
perspective? Yeah. From a platform engineer,
20:54
I often tell them the most important thing you
20:57
can figure out is what you're going to say no
20:59
to. Figure out the things that you're like, this
21:02
cluster is not for that. When I was at Disney,
21:06
Disney Plus had a platform internally. And the
21:09
thing we said no to was stateful workloads. And
21:12
it honestly was the best decision we had ever
21:14
made. Because anytime we wanted to move the cluster,
21:17
it was really easy for us to do. CVEs, anything
21:20
else, we're like, hey, if you need stateful stuff,
21:22
you go somewhere else. I'm sorry, this platform
21:24
is not for you. And and like if you can make
21:27
those decisions up front and you can say like,
21:29
OK, we have what we have, you know, five monitoring
21:33
systems and we know how we're going to upgrade
21:35
and we know how to control like all that stuff's
21:37
like just getting it set up for day one. You
21:39
deploy it, it can go to production. And at some
21:41
point you're going to say, OK, what what is it
21:43
that I have to say no to? And if you can figure
21:45
that out for day two, you'll be in a much better
21:47
place for day three, because at some point you
21:50
are going to have to upgrade that cluster multiple
21:52
times and you're going to have to at some again
21:54
talk to. teams that you don't know who's running
21:57
this workload and they don't care about it anymore
21:59
they're off to something new right and that applies
22:02
to any workload in the in the era of everyone
22:05
wants to be an AI company or at least everyone
22:07
wants to run AI workloads that becomes harder
22:10
because AI training is very stateful and needs
22:14
a lot of data and you need to figure out all
22:17
of your storage your networking performance tuning
22:22
And as soon as you get all that, you're like,
22:23
actually, this is going to run for a week at
22:25
a time to do some sort of like fine tuning or
22:28
training or something like that. And like now
22:29
I need to figure out how to snapshot it. Right.
22:31
And I was like, OK. And these are these aren't
22:33
new problems. These are problems that HPC environments
22:35
have had for a very long time. And so a lot of
22:39
what Kubernetes was built for, built for stateless
22:42
web applications originally, is now running head
22:46
head into this like job management and long running
22:49
jobs that it just it. isn't well suited for at
22:52
like the lower levels and and that's why people
22:56
still use a lot of Slurm and they still use a
22:58
lot of custom tooling for hpc and like strictly
23:02
hpc environments and in kubernetes it's getting
23:05
there at some point but i don't know i don't
23:09
see that It's not a perfect fit. You're going
23:11
to have to layer on some other abstraction to
23:13
get the features you want because Kubernetes
23:15
at the API level is like you can plug in whatever
23:18
you want and then go build it. And they're kind
23:21
of done putting in like core functionality into
23:23
the API. So given Kubernetes adoption, given
23:28
saying no to things, what's a practice outside
23:32
of that you wish more teams would have adopted
23:34
before they needed it? As far as like the structure
23:37
of their, from a platform engineering perspective,
23:39
as far as structuring for their team. It's something
23:42
that I don't think the platform teams can't give
23:44
themselves, but I wish that the organization
23:46
would give the platform team, the agency to have
23:50
like some financial power to say, to control
23:53
what's going on. Right. So if I look at like
23:55
SRE, SRE, when it came out, the Google book and
23:58
all this stuff, they're like, Oh, this is, this
23:59
is how you do SRE. This is, this is what S and
24:01
everyone's like, Oh, well it's like engineering
24:03
and production and uptime and all this stuff.
24:06
It's all true, but also the thing I took out
24:08
from that book was the thing that SRE at Google
24:11
did that no other SRE team ever did that I saw
24:15
was the SRE teams had the ability to say no.
24:18
The SRE team could say, your application sucks.
24:22
It has too much technical debt. I'm not going
24:24
to maintain it. And they could walk away. And
24:27
then the team was on their own to like bring
24:29
it up to a level of, yeah, we have to be able
24:32
to fix this now. And that was only because the
24:34
SRE organization had that ability, right? The
24:38
funding for SRE was taken out of engineering
24:42
budget from the application team. So if you wanted
24:44
an SRE on your team, you had to pay for it. And
24:48
then the SRE team could say, no, that's not worth
24:51
our time. Right. We have other places that we
24:54
should be spending our time. Most platforms are
24:58
not funded that way. Most platform teams don't
25:02
get their funding from the engineering team of
25:05
the application. Right. They're getting their
25:06
funding from infrastructure budgets or they're
25:10
getting their funding from maybe security teams.
25:13
And those are the places that are saying, hey,
25:15
I'm going to pay for this team. This is a 10
25:17
person team for building our platform. Why are
25:20
you building the platform? because we need a
25:22
platform. Tell me what you're trying to do. We're
25:24
trying to accelerate developers so they can have
25:26
less cognitive load about, developers shouldn't
25:29
learn Kubernetes, we should just subtract. I'm
25:30
like, well, what are you actually trying to do
25:32
here? And they're like, well, we want to centralize
25:34
things. Okay, what is the benefit of you centralizing
25:37
this golden path? And from my experience, the
25:40
main things that you want to centralize are security,
25:43
like being able to audit your systems and say
25:46
like, hey, is this going to be a problem for
25:48
us? If there's a CVE, you need to be able to
25:50
patch it quickly. If there's a leak or a hack
25:53
or something, you need to be able to audit your
25:55
systems and understand what's going on. So things
25:56
like logging, monitoring, things like just like
26:00
SBOMs and like security scanning, that's something
26:03
you need to centralize. And if you're in a cloud
26:06
environment, you need to centralize your costs,
26:09
right? Because like, that's just where are we
26:11
spending more money than we should be? That's
26:14
like the basics of a platform is usually those.
26:16
two things to me if you don't need to build up
26:19
the golden path for all this engineering stuff
26:21
if you just give them hey everyone now gets automatic
26:23
data dog logging here you go and everyone automatically
26:27
is in these AWS accounts and we're going to take
26:28
care of the how the costs bubble up into our
26:31
orgs so we can see that at central place and
26:33
everyone's going to get you have to have this
26:35
s -bomb in place in order to ship to production
26:38
right that's the basics of like a platform team
26:40
and I wish more platform teams or more organizations
26:43
would understand you don't want to centralize
26:46
everything. There's actually not a lot of value
26:48
in putting everything in the same centralized
26:51
tool, because again, that's just going to slow
26:53
you down at some point. Most teams are going
26:55
to say, we got to centralize CI/CD, because back
26:58
in the day, Jenkins was the way to go, and it
27:00
was really hard to set up Jenkins servers, and
27:02
so you had to set up one. And that one became
27:05
so critical to how everything worked in your
27:07
company that no one could upgrade it. And so
27:10
then you had this really old Jenkins server until
27:11
the next team said, you know what? I'm going
27:14
to send up my own Jenkins server because the
27:16
other one sucks. And then you had this like battle
27:18
of Jenkins servers in most companies. And at
27:20
some point, everyone's like, I like the new Jenkins
27:22
server. It has the new UI or it does the new,
27:24
you know, whatever, right? Like it's we want
27:26
to move to the new CI/CD. It's like, well, like.
27:29
Yeah, it's a cost to maintain those sorts of
27:32
things. And it is valuable to provide an optional,
27:35
hey, you can use this template that we have that
27:37
allows you to do something faster. But mandating
27:40
centralization through some sort of platform
27:42
team ends up being a hindrance. And most organizations
27:47
that I talk to and platform teams that I talk
27:49
to, they don't understand that. They just say,
27:50
we just have to centralize everything. We need
27:52
everything to run on Kubernetes. I'm like, no,
27:54
you don't. Trust me, you don't want everything.
27:56
And also like, this this keep this notion of
27:59
like you need kubernetes clusters to do anything
28:02
is that was like before AI right now we need
28:06
AI to do anything like before anyone's gonna
28:07
write any code like well if claude's down i can't
28:09
write any code like what are you doing like do
28:11
you do you remember how this works like you you
28:13
can't is it worth your time i don't know you
28:16
have to decide that right like how slow are you
28:18
going to be without it should you take a break
28:19
for an hour do you need to go refill your tokens
28:21
whatever right but like Back in the day, like
28:23
we keep doing the same thing just with a different
28:26
technology, right? Because even like 10, 15 years
28:30
ago, you put a credit card on a cloud provider
28:32
and that accelerated what you could do because
28:35
it took too long to get a server in the data
28:36
center. And then later you're like, okay, well
28:39
now I'm going to, you know, my credit card is
28:41
going to go to something else, some other service.
28:42
So I don't have to run this thing anymore because
28:44
I just don't have the time to. And now credit
28:46
cards are on cloud. And that's how like the amount
28:48
of people I know that are paying like the $200,
28:50
you know, thing, just so they're like, I just
28:52
need to stay like up to not up to date. And even
28:56
it's just like, I need to be. As productive as
28:59
the next person, I'm like, no, like, what are
29:01
you doing? Like, let's just understand where
29:03
the problem actually lies. And yeah, we're going
29:05
to keep doing this. There's going to be a new
29:06
technology that's going to come out in five years.
29:08
People are going to put their credit card on
29:09
it and they're going to say, actually, I just,
29:10
I have to have this to work. And I always sort
29:13
of like reevaluate that. Like, what are you actually
29:15
trying to get done? And when does it actually
29:17
need to be done by? And then once we have that
29:19
conversation, we can say, okay. well, then let's
29:22
plan this, right? It's just like a budget again.
29:23
If I know my workload is going to scale to maximum
29:26
of this in the next 12 months, I can plan for
29:29
that. And people are so against planning today
29:32
that they're just like, I just need to spend
29:34
money to solve the problems. So do you think
29:36
that AI kind of causes people to do less critical
29:40
thinking on their own? And I mean, the Dora report
29:43
last year says that we see this AI productivity
29:47
gain. across the industry, but we're not able
29:50
to quantify it correctly or quantify it accurately
29:53
in the data analytics that we get out of our
29:57
actual number of commits or number of features
30:01
that are released or whatever. So do you think
30:03
that there is this inherent symbiotic relationship
30:06
that people have formed now with their AI as
30:08
far as engineers unable to be engineers without
30:12
their AI tooling? I think engineers have formed
30:16
too close of... relationships with a lot of their
30:19
tools. It's not an AI problem. Vim and Emacs,
30:22
this is a problem that people have religious
30:25
debates about for a very long time about pretty
30:27
much any tooling because they feel attached and
30:30
they identify with those things. And AI is another
30:34
version of that. I think the fact that AI is
30:37
changing so frequently that people are having
30:41
a hard time even even figuring out what their
30:44
identity is because people that were co -pilot
30:46
advocates last year right if you were out there
30:48
today saying like co -pilot's the way to go everyone's
30:51
gonna laugh at you like no way man why are you
30:53
you gotta be on clod like what are you doing
30:54
right and this was this was literally eight months
30:57
ago that people were just like this was the way
30:59
to go and in microsoft just like everything's
31:01
gonna be co -pilot and now it's like they kind
31:03
of lost the open AI you know deals and now they're
31:07
just like well now we're just gonna run whatever
31:08
AI llm we think is the right one and and so like
31:12
Everyone's changing their mind about this stuff.
31:13
I do think that the kind of the I guess the status
31:16
quo of how people are working today where I want
31:20
to review code more than write code is more common
31:23
in some of these tools. I don't think that's
31:25
and I do think AI has accelerated that but I
31:28
don't think it's like completely new. Right.
31:30
Like I always think of Ruby on Rails was amazing.
31:33
The first time I ran it, where I was just like,
31:34
oh yeah, make me a new website. And like, it
31:36
would template out like 20 folders and all this
31:39
stuff. And it's like, you just put in these eight
31:41
variables and we're ready to go. I'm like, that's
31:43
amazing. Right. But like, and I didn't review
31:45
the code. I didn't look at it. I just said, okay,
31:47
well, yeah, Ruby on Rails, ship it. Like, you
31:48
know what you're doing. And we're, we're kind
31:50
of, it's, it's that, but it's. a lot more flexible
31:53
right like we're just we're doing the same sort
31:54
of stuff and even like go back a long time right
31:56
compilers and compilers they just take your you
31:59
know take your high level cobalt and and make
32:01
it machine code and that's amazing right like
32:03
you can actually do that without needing to review
32:06
it and at some point people move more and more
32:08
to like how do this architect together where
32:10
where are the important things for me to have
32:13
more input versus just being able to know like
32:16
oh i know how these pieces should fit together
32:17
someone else can actually like make the glue
32:20
or the templates for me So do you think that
32:23
with the advent of what we have, Claude, Mythos,
32:25
and Project Glasswing, do you think that that
32:28
changes the prioritization now more towards cybersecurity
32:31
on four platform teams and DevOps teams alike?
32:34
Or do you think that it's status quo and we just
32:37
work from the outside in, make sure there's a
32:40
WAF, lock everything down, and make sure our
32:43
packages are up to date? And that's enough. I
32:45
love how optimistic you are that people are actually
32:48
using the WAFs and everything. I mean, because
32:50
I look at what people are shipping today. none
32:55
of that is in place right this is just like oh
32:57
yeah we're just gonna we're just gonna burn through
32:59
this as fast as we can and again and that's i
33:01
mean to some of the larger you know companies
33:04
that have more to lose they have the policies
33:06
they have the practices they're gonna say like
33:07
actually no you can't ship that to production
33:09
without these things and engineers and people
33:12
are pretty smart they're gonna try to work around
33:13
it and figure out how can i get my thing done
33:16
to be successful my way everyone always tries
33:19
to work around some of that stuff but at the
33:20
end of the day like there is going to be some
33:22
controls on okay Where is this actually going?
33:25
What does it have access to? I think all those
33:28
attacks are going to be more sophisticated, but
33:29
I don't see it being measured yet. Like, S -bombs
33:34
are measured, right? Security is always reactionary
33:36
in that, like, oh, there's no CVEs in my architecture.
33:39
I'm like, yeah, there are. It's like, no one
33:41
knows about them yet, right? And I think that
33:42
that is going to be, until we can measure it,
33:45
we can't manage it. And so we have to flip this
33:48
to be a little more proactive. And WAFs are one
33:52
way of doing that. We're just going to have a
33:54
firewall, we're going to say no DDoS, whatever.
33:55
We're going to put some things in place, but
33:57
once your application's online, it's fair game
34:01
for anyone to start poking at it. And that is
34:05
still, I see, who was it just last week? Calendly?
34:09
They closed source their software. Actually,
34:12
we don't want people to see our code. We now
34:15
need to close source it because being open source
34:18
is a vector for... an AI bot to read our code
34:21
and say, oh, I see exactly where this is. At
34:24
the end of the day, it's probably going to be
34:25
some subscription and it's going to be tied to
34:27
some level of insurance, right? There's cybersecurity
34:31
insurance and some of those cybersecurity insurances
34:34
are going to say, hey, you have to run this mythos
34:37
on your code base once a month or something like
34:39
that. There's going to be some requirements there
34:41
just to say, we did what we could. But at the
34:45
end of the day, you still have to... if you don't
34:48
have a way for customers to access your stuff,
34:50
it's going to be really hard to sell something.
34:53
No, that's for sure. Okay. So closing out, what's
34:55
one piece of advice for platform or SRE people
34:58
trying to stay sane this year with everything
35:01
that's happening, all of the new features, all
35:03
of the new models, how do we stay sane? I would
35:07
say find the things that you're interested in,
35:10
right? Like don't just keep throwing YAML at
35:14
the wall, right? Like, like you actually like,
35:17
Figure out, oh, you know what? I really like
35:18
networking. I think storage is cool. I want,
35:21
you know, I want to understand some level of
35:23
something like find the things that you're passionate
35:25
about and that you think, oh, that's interesting.
35:28
Whether it's a career path or not, right? Like
35:31
whether it's like, you don't have to plan out.
35:33
I think this is going to be big, so I'm going
35:35
to go into it, right? Let's just like figure
35:37
out for you. Like when I first started Kubernetes,
35:39
it was not the winner, right? Like this was,
35:41
this was back in like Mezos was the thing that
35:44
everyone ran. And when I was like, oh. this kubernetes
35:46
thing is kind of cool and why is it different
35:48
than what Mesos offers or what Slurm offers or
35:51
and and just knowing that like it doesn't always
35:55
work out i have tons of projects that were just
35:57
throw away that like i learned a skill deep enough
36:00
that i don't need it anymore but it doesn't matter
36:02
that it didn't pay out for me like i i had fun
36:06
i learned the thing and in having joy in what
36:10
you do and is more important than like having
36:14
some form of like financial success, right? Like
36:18
the happiness, the joy, the being able to like
36:20
still be yourself is a lot harder. And platform
36:24
teams have a lot of breadth of things that they're
36:26
going to look at. SRE teams, you're going to
36:28
burn out no matter what. Like there's, oh, if
36:30
you are on call, I don't care where you're on
36:32
call for, that level of pressure is always going
36:35
to cause some level of burnout. So talk to your
36:38
manager early about, hey, how can I make sure
36:40
that like this is sustainable at a team level?
36:43
Like, hey, we know we're going to lose one person
36:46
a year to this. How do we handle on-call rotations
36:48
when we're short people? Those are all things
36:52
that you should be thinking about on making this
36:55
a little more enjoyable for yourself and not
36:57
just burning all your tokens and then going home
37:01
and just crashing. At some point, you're just
37:04
like, what do you have fun doing? And what do
37:05
you look forward to every week? If there's something
37:07
at work that you look forward to, you should
37:09
do more of that. Right. You should you should
37:11
say, like, actually, you know, I really like
37:12
this thing. When can I do that twice a week now?
37:15
Can I do it? Maybe, you know, like like dedicate
37:17
a week, a month or something. That's the important
37:20
stuff, because that's the stuff that will give
37:21
you the energy just to be a better. person in
37:24
general like you are more enjoyable to be around
37:26
when you like what you're doing in a lot of cases
37:28
and you like doing it with the same you know
37:30
right people and that's not always a choice everyone
37:32
can make it's not always a freedom that everyone
37:35
has but when you do have those things try to
37:38
do more of it and try to focus on it i 100 agree
37:41
with that justin where can people find your work
37:43
i'm on social media quite a bit i'm hanging out
37:47
a lot on blue sky um justin garrison.com is
37:49
my website i still blog there i've been blogging
37:51
for over 20 years and over that the entire 20
37:55
years i've i've averaged one blog post a month
37:57
awesome which is amazing right like i look back
38:00
at it when i migrated from a bunch of different
38:02
places into you know into my site and i was like
38:04
wow like i have a lot of articles like that's
38:06
pretty crazy and i just added them up and i'm
38:08
like actually that's that's amazing like i don't
38:10
blog all the time but i when i do i go into like
38:13
i'm gonna do three or four blog posts and I'll
38:16
put them all out there. And so I'm on average,
38:17
I'm about one a month. And then I also, in the
38:19
podcast Workaround Findouts, we're doing that
38:21
once a month as well. So I kind of have those
38:23
two things that I just try to keep up. And those
38:25
are where I'm interested in things and when I
38:27
want to talk to people and where I find joy.
38:29
Awesome. Well, I'll put links for all of that
38:32
in the show notes. And Justin, thank you so much
38:33
for coming on. Really appreciate it. Yeah. Thank
38:36
you. That was my conversation with Justin Garrison,
38:39
Field CTO at Sidero Labs. and co-host of Fork
38:44
Around and Find Out. The biggest takeaway for
38:46
me is that good platforms are defined as much
38:50
by what they refuse to do as what they support.
38:53
It is tempting to centralize everything. One
38:57
Kubernetes platform. One CI/CD system. One golden
39:02
path. One way every team is supposed to work.
39:06
And that can absolutely make things easier at
39:09
first. But eventually... The exceptions pile
39:13
up. The platform absorbs more responsibility.
39:16
And abstraction that was supposed to reduce cognitive
39:20
load becomes its own source of cognitive load.
39:24
I liked Justin's example from Disney +, where
39:27
the platform simply said no to stateful workloads.
39:31
That constraint made everything else easier to
39:35
operate, upgrade, move, and recover. That is
39:38
a useful way to think about platform engineering
39:40
in general. Do not start with how do we support
39:43
everything? Start with what problem are we actually
39:47
solving? What should be centralized? And what
39:50
should explicitly live somewhere else? The same
39:53
idea applies to Kubernetes and AI. Once we learn
39:58
a useful tool, we tend to want to use it everywhere.
40:01
Kubernetes becomes the answer to every workload.
40:05
AI becomes a requirement before somebody writes
40:09
a line of code. Sometimes that is useful. Sometimes
40:13
we are just replacing one dependency with another.
40:17
So my takeaway is to keep questioning the abstraction.
40:21
Know why you are using Kubernetes. Know why you
40:25
are centralizing something. Know what your golden
40:28
path is optimizing for. And know when the right
40:31
answer is simply this platform is not for that.
40:36
You can find Justin's work, his writing, and
40:39
fork around and find out through the links in
40:42
the show notes. Follow or subscribe to Ship It
40:45
Weekly wherever you listen and find previous
40:48
episodes at shipitweekly.fm. Thanks for listening,
40:52
and I'll see you later this week.
The thing that stuck with me most from this conversation with Justin Garrison is that mature platforms are defined just as much by what they refuse to support as what they enable.
That sounds backwards at first.
When teams build platforms, the natural instinct is to think about capability. How many workloads can we support? How many teams can use this? How many tools can we integrate? How much can we automate?
But every additional capability creates another responsibility.
Every exception adds complexity. Every supported workload type adds another path to maintain. Every abstraction creates another thing engineers have to understand.
At some point, the platform that was supposed to reduce cognitive load becomes another source of cognitive load.
That is the trap Justin described throughout this conversation.
Platform engineering is not about building the biggest possible platform. It is about creating the right boundaries.
One of the examples that stood out to me was Justin’s experience at Disney Plus, where the platform made a deliberate decision not to support stateful workloads.
At first glance, that sounds limiting.
A platform that says “no databases” or “no stateful applications” might feel like it is making developers’ lives harder. But that constraint created simplicity everywhere else. Cluster upgrades became easier. Ownership became clearer. Operational concerns were pushed to the teams and systems that were actually designed to handle them.
The platform became better because it knew what it was not.
That is a lesson that applies far beyond Kubernetes.
We have a tendency in technology to take something useful and expand its scope until it becomes the answer to every problem.
Kubernetes is a great example.
Kubernetes solved real problems. It gave teams a consistent API for deploying workloads, managing containers, and automating parts of infrastructure operations. It became the foundation for modern platform engineering for good reasons.
But somewhere along the way, Kubernetes also became the default answer for almost everything.
Need to run an application? Kubernetes.
Need internal tooling? Kubernetes.
Need batch processing? Kubernetes.
Need AI workloads? Kubernetes.
Need a simple service that could run on a VM for five years without anyone touching it? Probably Kubernetes.
Sometimes that is the right answer.
Sometimes it is just the tool we know.
Justin brought up a great comparison with Ansible. People still use Ansible for things it was never really designed to do because it became familiar. It was the tool they trusted, so every problem started looking like an Ansible problem.
The same thing happens with every successful technology.
We do not just adopt tools. We build identity around them.
That is why the conversation around Kubernetes, platform engineering, and even AI is so interesting right now. The challenge is not learning another tool. The challenge is knowing when the tool is actually helping.
The discussion around cloud versus on-prem infrastructure was another area where I thought Justin brought some useful nuance.
For years, the cloud conversation was often framed as if moving to cloud was automatically the mature decision. The assumption was that companies moved from servers to virtual machines, from virtual machines to cloud, and eventually everything became cloud native.
But reality is more complicated.
Cloud solved a lot of problems. It gave teams access to infrastructure faster, reduced the need for massive upfront investment, and created incredible flexibility.
But cloud also introduced new problems.
Cost management became a discipline of its own. FinOps teams exist because cloud spending can become complicated quickly. Autoscaling sounds simple until you realize scaling infrastructure often means scaling databases, moving data, managing networking, and handling operational complexity.
At some point, some companies look at their environments and decide that more control, predictability, or sovereignty matters more than unlimited flexibility.
That does not mean cloud failed.
It means architecture is about tradeoffs.
The right answer depends on the workload, the business, the regulatory environment, and the operational capabilities of the team.
That same idea applies to Kubernetes distributions.
Justin’s experience working on EKS Anywhere was interesting because it highlighted something people sometimes overlook. “Kubernetes anywhere” does not necessarily mean the same thing everywhere.
Managed Kubernetes offerings are optimized for their environments. Running Kubernetes on-prem is a different problem. Running Kubernetes at the edge is a different problem.
The API might look familiar, but the operational model underneath can be completely different.
That matters because portability is often oversold.
You can move Kubernetes workloads between environments, but the deeper assumptions around networking, storage, identity, observability, security, and operations do not magically disappear.
The YAML might move.
The architecture usually does not.
That leads directly into the platform engineering discussion.
Justin’s advice for platform teams was simple but probably one of the most important points of the conversation.
Figure out what you are willing to say no to.
That is the part many teams skip.
They start with what they want to build. A developer portal. A golden path. Self-service deployments. Internal tooling. Kubernetes clusters.
But before building any of that, the harder question is: what problem are we solving?
Are developers struggling because they do not have enough automation?
Or are they struggling because approvals take two weeks?
Is the deployment process slow because there is no platform?
Or because nobody owns the testing strategy?
Is Kubernetes complexity the problem?
Or did the organization choose a tool that created complexity it did not actually need?
The platform is not the outcome.
The outcome is enabling teams to deliver software safely and effectively.
The same thinking applies to centralization.
There are some things that make sense to centralize.
Security controls. Compliance visibility. Logging. Monitoring. Cost visibility. Supply chain security.
Those are areas where consistency and visibility provide real value.
But centralizing every developer workflow, every deployment pattern, and every engineering decision can create a bottleneck.
A platform should remove friction.
It should not become the new approval queue.
The AI conversation followed a similar pattern.
One of the interesting things about AI right now is that it is not introducing an entirely new human behavior. It is accelerating a very old one.
We have always become dependent on tools.
People built strong opinions around Vim versus Emacs. Developers became attached to languages and frameworks. Teams built entire processes around CI systems and deployment platforms.
AI is another version of that relationship, except the speed of change is much faster.
Six months ago, one tool was the answer. Today, another tool is the answer. Six months from now, something else will probably replace both.
That creates a challenge.
The risk is not using AI.
The risk is forgetting how to think without it.
Justin’s point about fundamentals really resonated here. The engineers who will benefit the most from AI are not the ones who know how to generate the most code.
They are the ones who know when the generated code is wrong.
A Kubernetes manifest can apply successfully and still be a terrible architecture.
A Terraform plan can complete successfully and still create a disaster waiting to happen.
A generated application can compile perfectly and still have the wrong security model, scalability assumptions, or operational behavior.
The syntax is becoming cheaper.
The judgment is not.
That changes what experienced engineers bring to the table.
The value moves higher.
Architecture. Tradeoffs. Failure modes. Security boundaries. Understanding systems deeply enough to recognize when something looks correct but is actually dangerous.
That is the part AI cannot replace easily because it requires context.
And that connects to Justin’s final advice about finding the parts of engineering you enjoy.
Technology changes constantly.
The tools that define our careers today might be completely different a decade from now. Kubernetes replaced other platforms. Cloud replaced other infrastructure models. AI is changing how we build software.
But curiosity lasts.
The engineers who stay successful are usually the ones who keep asking questions.
Why does this work?
What happens when it fails?
What assumption is this abstraction hiding?
What problem are we actually solving?
Those questions matter more than any specific tool.
Because every new technology eventually becomes someone else’s old technology.
The fundamentals are what carry forward.
Additional Links
Justin Garrison:
https://justingarrison.com
Justin Garrison on BlueSky:
https://bsky.app/profile/justingarrison.com
Sidero Labs:
https://www.siderolabs.com
Talos Linux:
https://www.talos.dev
Fork Around and Find Out podcast:
https://forkaroundandfindout.com
Kubernetes:
https://kubernetes.io
Amazon EKS:
https://aws.amazon.com/eks/
Amazon EKS Anywhere:
https://anywhere.eks.amazonaws.com
Kubernetes Cluster API:
https://cluster-api.sigs.k8s.io
Kubernetes The Hard Way:
https://github.com/kelseyhightower/kubernetes-the-hard-way
Google SRE Book:
https://sre.google/sre-book/table-of-contents/
FinOps Foundation:
https://www.finops.org