Hi, Chris. It's good to chat with you about DataOps. I think this is a topic that when I think of an expert of DataOps in the field, you're the first one that comes to mind, so it's awesome to see you. It's awesome to see you. Thank you. Yeah, of course. I guess maybe give a quick intro for people who don't know who you are. So I'm Chris Berg. I run a company called Data Kitchen here in Massachusetts. And so I'm an engineer by training, sort of built software for a bunch of years. I've had different companies like MIT Lincoln Laboratory, NASA, some Internet startups. And then I got the bright idea, like in 2005, that I should do data full-time. And so I started to manage a team of what we now call data engineers and data scientists and people who did data visualization. And sort of my life really sucked. Things were breaking. We could never go fast enough for our customers. And, you know, I fired a bunch of people I probably shouldn't. I blamed people for problems when, in fact, it really was a management problem on how I managed that team. And so the question that I've been working on for the past almost 20 years is how do you get a team of people, a team of technical people doing data to deliver things quickly, new insight quickly, how to deliver data and dashboards with really high quality, and then sort of like how to live, like how to not have a painful life. And so I think it's great that you're teaching a course on data engineering, but we did a survey with Data.World two years ago of 700 data engineers, and something like 78% wanted their job to come with a therapist. And, you know, I guess maybe I'm cynical, but I was like, oh, that's a little low. And I think it is a very difficult career to have, sort of caught between the rock and hard place of sort of insane customer demands and data and infrastructure that's just hard to live with. And so, you know, I came to the sort of ideas of DataOps sort of from that background. Interesting. I guess for the learners, what is DataOps? How would you describe it? So think of it as a method to make yourself or your team deliver insight to your customers where they can trust the data and the insight, and then you can make changes to that insight very, very quickly with low risk. That's really the idea is can you build a factory and push out new data sets and new insights? That's perfectly, perfectly good. And change things in the factory as fast as you want. It seems to take inspiration and lead, it sounds like. Yeah, I think that's one of the things that surprised me when I started to do data analytics, right? I was a software guy. I'm like, oh, I'm great. This data thing will be easy. And then you really are running a manufacturing line in a lot of ways. And if you think about a data set coming in and being put in a bucket and a database in multiple levels and then being used in a model and a visualization and then governance, there's sort of each step along the way. You could kind of think of it as a manufacturing station. And then, you know, you don't have one, right? You have lots of them. And so how do you get that manufacturing line to put out, you know, Toyotas and not sort of AMC Pacers from the 70s? And that's a hard problem. That wasn't really in my software background, like how to run a manufacturing line. And then the other weird part of it is that people are always going to ask you for new things. You know, your job's never done. And the best way to deliver new insight to a customer is sort of get it 80% done, get some feedback, and then iterate on it. And you don't spend months building something when you can get some feedback early. Those are two very opposite things. Build a really good assembly line, but change the assembly line really quickly. And those are hard to do. Right. If you go back to places like Toyota, they were able to swap out manufacturing in their assembly line. I was thinking days, not months, which is sort of the old mass manufacturing way of doing stuff, right? Yeah, there is. And I think that's how did they do it. They didn't do it by, like, buying fancy robots. I mean, they have robots, right? But the solution to that is really about how do you think about what your job is and what sort of things optimize. And so a lot of teams try to optimize on getting things done. Like, I have to get my work done. And the definition of done is really interesting, right? I've done my work. I've got my SQL. I'm done. And does it matter if someone used it? Does it matter if it breaks next week? I don't know. It's not my job. Right. And so, really, it comes down to it's about Toyota, but Toyota also focuses on the customer and how do you make customers successful. And, unfortunately, a lot of people in data and analytics are not happy with their โ the customers are not happy with data and analytic teams. And that's the dirty little secret is that most projects in data and analytics fail. And to me, it's not because we don't have cool tech, and it's not because the people aren't smart. It's because the systems that people work in are kind of flawed. And so these principles that come from software and manufacturing, I think, really apply. And they're not hard. They're not rocket science. They're just ideas that are laying there that you can pick up and use. Absolutely. So if somebody is new to data ops, how should they think about this in their job? Well, I think about the thing that you do. I'm building โ I do some SQL. Then think about the thing around it. So number one is I build some SQL. How do I build a test or a validation to prove that that works? Right. And I get some new data in it. How do I know? Now, and why do you want to do that? Right, because tests are sort of the gift that you give to your future self. If you don't do that, you end up owning your piece of code forever. And you have to โ and if something goes wrong, they call you, maybe on your vacation. Or you get stuck owning more and more and more code. And it's really tempting when you first start to be a hero to, like, I'm a hero. I wrote all this stuff. I keep writing it. I respond really fast to my customers. And you go into full hero mode. And what happens is, over time, you get burnt out. And you get unhappy. And you want a therapist. And so in one lesson of DataOps is be a hero maybe two, three percent of the time. But the rest of the time, build a system so you don't have to be a hero. Right? And that's what it's about, building a system next to your code. Testing, observability, deployment. It's all things that are sort of the environment around that squishy little cool code that you're writing that can make your life tractable. And then let's talk about the DataOps manifesto. So what inspired that? Like seven, eight years ago, we went to a conference, and no one knew what we were talking about. And I explained DataOps, and they're like, what's that? And it was like, oh, man, this sucks. No one really gets what we're talking about. And so we had to โ so I wrote the โ on the plane, I wrote the first version of the manifesto. I mailed it around to some people, got some feedback. I wrote the Wikipedia article. And we've had to write a lot over the years to talk about what DataOps is. We have a free certification. You can listen to me for three hours on our website. But really it's this idea that building systems of work is really important and sort of solving your team for kind of a couple problems. Like if you look at it, you hire a new 23-year-old on your team. They're new to data engineering. Maybe they have a CS degree. What do you want to do with that person in their first two weeks? I think a good organization says, number one, if something goes wrong, they should be able to find the source of the problem. Is it data? Is it the integrated data? Is it the model? Is it the biz? Like where's the problem? Or number two, like could they write some small โ could they tweak something? Like anything, fix a bug, and get that into production quickly. And those things are really hard in a lot of data engineering teams. You have to kind of learn things, and it almost becomes an apprentice. And so, like I think of what data ops is. You've got two 23-year-old problems, how fast the 23-year-old can find a problem, how fast the 23-year-old can fix and deploy a problem into production with low risk. And then there's also probably a 46-year-old problem, which is like how do you measure your team. And to me data ops is about those sort of problems, which really they don't really have to do with data engineering. It has to do with how you sort of manage a group of data engineers working together. And that's sort of where I come from. And there is a whole engineering task. I mean writing a good data quality validation test is important. Deployment speed is important. Doing DevOps, managing environments, test data, all those things are really important. So there is skill in that. One question I'm sure that learners will have is how does data ops differ from DevOps? So I think in a lot of ways they're the same idea. They're all like how do you get a group of technically nerdy people to deliver things quickly with high quality. And you certainly can do it badly both ways. And so what DevOps idea is is that you should deliver your work in small chunks. Like don't work for months on something because if you do that you're most likely going to miss what your customer wants. And so the main idea that's similar between DevOps and data ops is if your customer asks you for 10 things and it's going to take you three months, deliver one of those things in a week. Don't wait and get all 10 things done because most likely what's going to happen is you're going to get four things that they want. They're not going to want six things. And then they're going to say, no, I want three more. And so you have this waste, right, of doing all this work that didn't really matter. And so I think the idea of DevOps is how do you manage short cycle times of delivery and how do you automate that. And I think there is a sort of observability error reduction in production in DevOps, but I think with data ops or whatever word you want to put on it, it's the same idea, right, is that run things with low errors and change it quickly and measure your work. And they're very much the same ideas. So whether you call it lean or whether you call it TQM or whether you call it DevOps or data ops, they're all the same things. How do you get nerdy people to be productive? And to do that, you actually have to do some things. It's not just words. It's actually work. You have to build some stuff. You have to build the thing next to your environment. You have to build the manufacturing line. You have to build the machine that makes the machine. Well, thanks. This has been very educational. I guess to cap it off, is there anything that you wish you would have known about data ops that you could tell your younger self, maybe the learners for this course? I think the two things is hope and heroism. It's like, number one, what I said before, don't be a hero. You're not doing yourself any favors. And then the second part is don't hope that things will work. Maybe it sounds paranoid, but don't trust your data providers. Don't trust your servers. Measure and prove that things work, both of which I've had so many problems with in my career, just hoping things will work or trying to be a hero. And you can build systems around that. So you can have a little hope, have a little heroism, but it's better to build a high-functioning organization that can do that. And if you do, your just life's going to be a lot better. And you don't want to end up in a situation where you're like that 80% of the people we surveyed and said, my job sucks, you know, I need a therapist, I'm so stressed. Awesome. Well, thanks for your time, Chris. This was very helpful for the learners. So thank you.