So for this lesson, we have none other than Wes McKinney to join me for a chat today. So looking forward to discussing pandas as well as I think other data related topics, Wes. But for the learners who don't know who you are, do you want to give a quick intro? Sure. My name is Wes McKinney. I'm the original creator of the pandas project for Python all the way back in 2008. Open sourced the project at the end of 2009 and built an open source community around it. I wrote the book Python for Data Analysis, which is now in its third edition. And a lot of people use that book to learn pandas and it's become a staple reference textbook for the ecosystem. I've worked on a bunch of other open source projects like Apache Arrow and Ibis, which is another Python project. And I've I've been a I would call it an open source entrepreneur. So I've started a number of companies, Datapad, Ursa Labs, Voltron Data, basically building software that's connected with this new ecosystem of open source data processing and data science tools. And, yeah, I also more recently have I invest in data companies and just generally interested in the growth and development of open source data science and making people making it easier for people to to work with data. That's super awesome. Yeah. And I think I had the original pandas book back from I think it was 2010 or 2011 or something like that that it came out. I can't remember the exact date, but yeah, that's the classic. So I guess for the learners out there who are new to pandas, what exactly is pandas? Well, pandas is a data like a data management or data manipulation toolkit. So it helps you load and manipulate data sets in memory and Python. So if you have a CSV file or an Excel file or a database, one of the first steps that you to do any kind of data data work in Python is that you need to load the data into memory so you can manipulate it as a table or spreadsheet like object. And so pandas provides you something called a data frame, which is a tabular data structure that that has an API for doing spreadsheet like operations, column operations, data cleaning operations, and then a lot of functions that help you load data out of data files and external data sources into into pandas data frames. There's obviously a lot of other stuff in the project. You can use it to generate plots and to create nicely formatted tables for inserting into publications. It's been a very popular tool for doing interactive exploratory data analysis in Python, and it's often the first stop on the way to using some other toolkit for for doing statistics or machine learning in Python. So if you statsmodels or scikit-learn or TensorFlow or PyTorch, often people are using pandas to load the data, do some data cleaning and data data manipulation or merging data sets together. And then they export they they hand off the data from pandas and they use it in one of those other those other libraries. So pandas isn't trying to be a an A.I. or machine learning toolkit. It's just helping you work with data. Yeah, I remember the first time I encountered it, I think it was 2012 or something like that. And I did something in feature engineering for machine learning at the time. And seeing Java, that seemed like the libraries back then were just pretty janky and very hard to deal with. But then I got pandas and it felt like all of a sudden my data was the equivalent of clay. I could mold it in any possible shape. So what inspired you to create pandas, though? I mean, what I'm sorry. Yeah, I was going to say it's it's it's it's amazing like how how much time is spent just like just getting access to data and loading, being able to read a CSV file and have it work reliably is is a big deal for for for a lot of people. So it's a big, big deal. Yeah, I guess kind of going back to, you know, circa 2008. Like what what inspires you to create pandas? Well, I was in I was in my first job. I was working for a quantitative hedge fund. It was during the great, great financial crisis. And we were under a lot of pressure to do a lot of exploratory data analysis and research to be able to make decisions faster. And to like many of the tools that I was using, there was some people using MATLAB. People were using R. There's a lot of Java and C++. But I wasn't really very satisfied with any of the a lot of SQL, too. We were using Microsoft SQL Server and I discovered Python and I found, wow, that Python is really productive, nice programming language to to program in. And but it was missing a missing a data toolkit. So I started doing some of my research, some of my data analysis work in Python, and then started building tools for myself just so that I could just so that I could do my work. And I found that I really enjoyed building tools. And I started socializing my little data toolkit with my colleagues. I didn't even have a name back then. It wasn't called Pandas or anything like that. And I found that my colleagues really liked using it. And eventually it got to a critical mass and where it made sense to like, hey, this might be useful outside of this company. Like this could really this could have an impact in the open source, the open source world. What did the name come from? I was working with a lot of econometricians who were always talking about panel data. And so that was so so that that partially motivated the Panda. And I was also trying to find some kind of an acronym or like a mishmash of letters and words that would evoke the idea of like Python data analysis. And so in kind of writing all those things down, I was like, oh, there's pandas occurring here. So that's that's how the name, you know, got into our mind. OK, this is also one of the biggest mysteries, biggest questions I've had. So that's interesting. Yeah. And it feels like the import pandas as pd has pretty much been burned into every data person's brain by this point. But obviously pandas wasn't always popular. What do you think accounted for the rise and success of pandas? I mean, there was there and there were a number of factors. I think in the 2012, 2013 era, there was a there was an enormous need for people, people with data science skills. And so programming languages where doing data science work was really accessible was a was an untapped opportunity to to expand access to to data science. And doing it in open source software meant that there was no paywall. You didn't have to buy by software. You didn't have to buy MATLAB to get to get access. And so I think the combination of like there being this great need for data skills and data science, a lot of new tech companies who were who were data focused were comfortable using Python and open source software. And they were looking for a data toolkit. And so they found they found pandas. Having my book come out, Python for Data Analysis in 2012, also helped because people could pick up that book and say, OK, I can read this book and learn how to do basic data work in in Python. And yeah, so it was a combination of like the right place and the right time. There being this great need for great need for data scientists. And, you know, I think people starting to be more forward forward thinking about using adopting open source, open source software for for data work is as composed as compared with with with with open with closed source software. Oh, yeah. Yeah, that was a big one. Like, I think, you know, I kind of consider pandas to be. It felt like that was sort of the gateway drug to data science in a lot of ways where data science was taking off. But it just made it so much easier because I was running a Python meetup, I think, 2012, 2013 or something like that. And the discussions back then in Python were mainly about web development. Like data wasn't really take off. It was about the inflection point. And I think within a year, all the talks at the meetup were just machine learning and data. And, yeah, pandas used to prominently. Yeah. Yeah, because this was like the like the early 2010s was the era of like the original strata conferences. So you had like like big data was everywhere. Everyone, every company in America and in the world was trying to figure out how to do big data, like how to do how to become a data company. And and so, yeah, I think there there was like there was definitely like a gap that that that needed to be needed to be filled in terms of tools. And, yeah, I think, you know, Python, for a variety of reasons, became the mainstream language for for data science and machine learning and in particular. So having Google and Google and Facebook choose Python as the as the language for for their machine learning frameworks, TensorFlow and PyTorch was a big deal in terms of cementing, cementing Python as the language to learn and use. Oh, yeah. It's been amazing to see the success of Python. So one part of the course that the learners are going to to explore is when is something like pandas a good fit? And when might they want to try a different approach? Pandas is a good fit when your data is tabular or rectangular, or if you have multiple multiple data sets that are that are tabular that you need to to integrate together. So it it fits that like spreadsheet like or database like, you know, interface of working with a table with, you know, column names and data in those columns. You can use pandas to create multidimensional pivot tables and more complex, complex objects. But it's strong suit is really that really that that tabular, rectangular, rectangular data. Interesting. When when might you opt for for maybe a different approach to working with data? If data sets are are really large, like they're too large to fit into memory or either they strain the limits of what you can work with on your machine or your laptop, you might consider using a you might consider using some type of some type of database or there's another there's other data frame libraries out there. There's a new new Python project called Polars, which you could you could try using, which is can is built for working with much larger data sets in in in Python. But most people have pretty modest sized data sets. And so that, you know, for most pandas users, that's rarely that's rarely an issue. But when people truly have like these massive hundreds of millions of rows or billions of rows, then that's where that starts to can start to become start to become an issue. And yeah, and there's yeah, there's plenty of users who are working with mostly machine learning data sets that are like clean and like nice, you know, multidimensional arrays, multidimensional arrays stored on disk. They might be using something like xarray or another multidimensional array framework, which is built on top of NumPy. And there is intended for use with PyTorch or TensorFlow. And so there may be instances where using pandas is not not necessary at all because you can load your machine learning data sets and be and be off to the off to the races. Also, I think that's really useful for the learners. I guess a lot of the people taking this course, I'm guessing, are going to be aspiring data practitioners. Do you have any advice for people looking to get into the field of data? I would say, like, in addition to learning the data manipulation, loading, cleaning, integrating, merging, like the basics of data wrangling. I think the becoming proficient at some of the exploratory data analysis skills like plotting visualization, looking at the data, getting good at using the, you know, the Jupyter notebook and different features of of the Jupyter notebook. I think a lot of what helps helps people be more productive, especially when they're just getting started, is is being able to look at the data and have like more feedback and more be able to have like a more iterative cycle with their with their analysis, as opposed to, you know, working in a editor editor tab of VS code and writing, writing a panda script and then running that panda script on the command line. So developing a more interactive, iterative approach that's based around looking at the data a lot and visualizing it can help you catch your mistakes faster and and ultimately master the skills that you that you need more quickly. That's great advice. And kind of closing out, where do you see the fields of data? I guess the broadest sense and maybe data engineering specifically, where do you where you see these things going over the next few years? I mean, I think that we'll continue to see evolution of the tools and the frameworks for for for Python in particular. I think Python will continue to be a mainstream, mainstream programming language for for for data work. I think the role of AI assistants like Copilot and ChatGPT and LLMs in the loop will become more and more natural and something that's there when you need it, but doesn't doesn't get in your way because, you know ChatGPT and other LLMs like they know an enormous amount about pandas because of all the content that's available on the Internet that has to do with pandas. And so I think that can be a really powerful tool to help you figure out problems or fix errors in your code or come up with new ideas for for how to work with your data. And so I think that's just yet another tool in the toolbox that's going to increase accessibility and make people a lot more a lot more productive. So we can automate more of the boring work and spend more of our energy on the more creative and value add kind of model development and and and data analysis. That's awesome. Great perspective. Thank you very much for your time. It's been great chatting with you. And thanks again for your contributions to the field. I think that your pandas is definitely it's a mainstay. So thanks for that and everything else you've done too. Yeah.