Hey, what's up, Juan? How are you doing? Great, Joe. How are you doing? Good, good. So when I think of graphs and graph databases, you're the first person that comes to mind. So I figured let's have you educate the learners on graphs. So I guess for the learners who don't know who you are, do you want to give a quick intro by yourself? Yes. So I am Juan Sequeda. I am the principal scientist and the head of the AI lab at data.world. We're an enterprise data catalog, but on top of a knowledge graph architecture. And graphs and knowledge graphs have been what I've dedicated my entire life, career at. So yeah, I'm here to talk about all things graphs. And I can talk about this for hours, if not days. So I think we've chatted about it at least for hours before. Yeah, so let's start from square one. So what's a graph? All right. So a graph is a data structure which consists of nodes and edges. And if you think about it mathematically or actually from set perspective, you have a unary predicate or like a unary. Imagine a set only has one thing. So just imagine like a table with only have one column, a unary. And then you have binaries, which is going to be basically imagine a table has two columns. So that one that has one column is going to be the node. And the one that has two columns is going to be the edges. And so the bubbles, the dots that you draw and the lines are the edges, which are going to be the relationships between those. And that's a basic structure. Is this kind of like parent and child relationships? Is that a way to think about this? Now we're getting into semantics and stuff. But yes, that's why you can see that the semantics and the knowledge part lives on top of the graphs. The moment that you start saying, hey, that edge, that relationship, there's a meaning to it. Which is more than just saying A is connected to B. So let's back up here. So we have nodes and edges and then more nodes. Is that a way to think about it? So you've got a node and an edge and then another node here. And what do these represent? So technically a graph could be just... You could have a graph which has only one node. That's it. I can have a graph that has just a bunch of dots. That's it. Then I can say there's a graph and then there's an edge that connects these things, a line, right? That's that relationship between two dots. And then this is how things start to grow. And then you can start seeing the characteristics of a graph. Is how many edges does it have? How are there clicks within a graph? What are the paths between different nodes in their graph? And those are the types of things that you'd be asking around just a graph structure. Interesting. And so let's get into this a bit then. So there's graphs. And then for the learners in this course, we'll be using a graph database. What's... How do these things relate to each other? So let's go back into what is a database. So a database, right, is a collection of data. I remember my first courses on database. Somebody pulled out... Remember those yellow books? Is this a database? It is. You're collecting data in there, right? And then you can do... And then... So that's the database is a collection of data. So then we start thinking about a database management system, a DBMS. So database management system is a way how, again, I manage the data that's inside that collection. And inside... And then you have different types of management systems for databases. The traditional one we hear about are relational databases. And that's where you are using a particular model for the data called the relational model, which is basically our tables that we go do. Then you have a graph database, right, which is really a graph database management system, where it's a system that has been designed to go manage data in the form of a graph. And by management of data, it's usually all the CRUD operations, right? So it'll create, read, update, and delete anything inside of that collection of data. So graph databases. Then you have relational databases. And then you have this whole era of NoSQL databases that ended up being that people were basically saying they're not just the SQL relational ones. So that's where graphs would come up. You have the document databases, which are really the JSON. I mean, heck, there used to be that people would call XML databases. Back in the 80s, we had object-oriented databases. Where a lot of the XML and kind of document stuff came out. And a lot of the object-oriented databases from the 80s has a lot of the graphs and stuff. So that's a little bit of where all this database has come from. Okay, so we could use a database and, I guess, query a graph then. Is that what you're saying? Yeah. So that's part of the thing is that one thing is that you have the storage, right? A database stores things, but you also need to be able to go access them. So that's where you have query languages, right? So you have a database query language. Databases have a manipulation language, which is how you do inserts, updates, deletes. And then they also have usually the schema, right? So basically in SQL, it's going to be your create table statements you want to go do. So if you think about a relational databases, you have your query languages, your select from, right? And then you have your selects, where you're querying from. You have how you can insert data, delete data, update data. And then you have all your create table statements. That's how you create the database. You have the equivalent of those things also inside of other graph databases. Got it. And so let's walk through a couple models here. The one that comes to mind is something that I've seen called RDF. And there might be another one called a property graph. But what's the difference between these two? Okay, so there are two different types of graph models. So again, just to the analogy, people get SQL all the time, right? So SQL is just your tables, and you have a standard language called SQL. And that's actually standardized by ISO. So it's one of the organizations that's been standardizing this for 40 plus years or more. So then what happens, and by the way, the relational database, the model comes from the database theory by Codd, right? That's where all that comes from the 70s. So then you have that same type of theory and stuff that has developed from in the graph world to two perspectives. One is called RDF, which stands for the Resource Description Framework. And the other one is called property graph, which is really kind of originally was popularized a lot by the folks by Neo4j very early on, like 10, almost 15 years ago, I would say. So let's go down those two routes. The RDF route, the RDF actually comes from the web. So it comes from the web community. So what happens when the World Wide Web Consortium starts out as kind of the standardization for the web, the web is all about having, it's a graph of documents, right? These HTML pages, documents that can point to another document. And what really happened was that the original vision of the web by Tim Berners-Lee was it's not just about documents pointing to other documents. I really want data to point to other data and be able to describe what this data means. So RDF was the manifestation of the standard for data, while HTML is a standard for text. So you can imagine text, web pages have HTML, data has RDF. So because you're following the graph, you're following the web, which is you have a page, links to another page, you're basically creating a triple. That's where the triples come from. So the RDF is a model based on triples, which is subject, predicate, and object. So it's predicate and object, okay. Going back to the graph model, node, edge, node. Subject, predicate, object. And just how we talked in English, you have subject, verb, object. So these things are all very related too. So that's what the RDF model comes in. And it comes, it actually was originally kind of proposed, presented back in the 90s, I think like 98, actually. And it was really a way to start, the first uses of it was, how do I describe web things on a web page? So it was basically adding common, just metadata annotations on a document. And it was all just triples, all just part of a graph. So you would say, hey, this document was written by Joe Reis. Was written, this document was written on, the date that it was written was January 1st, right. So you're adding all the graph things. What happened is that, that was made for the web, kind of for annotations. And later on, people said, oh, now we're creating more and more of this data on the web. We need to have a way to go store it. So they kind of, the database was an after the fact. So then database systems came to go manage RDF. And you start seeing languages that were standardized by W3C called SPARQL. So SPARQL is the graph query language for RDF. And then we started seeing how to go manage more of the semantics, because these links have a meaning. I want to add the semantics, the meaning, the modeling around this. So just so you're saying kind of taxonomies and stuff, those ended up being part of the standard. And things like RDFS called RDF schema, and then things like OWL, which is the web ontology language. We can go more into that stuff. So there's that whole standard over there on that kind of triple based model. The property graph is something that it's, again, also nodes and edges. But what you can think about them is that you have key value pairs that can be associated to a node or associated to an edge. And that's really the other model they have. And I think the way the property graph started was more, was less about thinking about, I want to share data on the web. It was more about, oh, I want to create graphy applications. So edges are first class citizens, and I want to be able to store things around the nodes and the edges. And so they were a database first mentality. I think originally, Neo4j didn't even have a query language. It was all through APIs. And then they created a language called Cypher. And then other different database systems followed that property graph model, which is key value pairs on nodes and edges. And then I think there was this big movement to kind of create a standard for a property, for a language. And what Neo4j did was kind of open source a language called OpenCypher. And then in the last couple of years, a lot of folks got together to say, hey, kind of different visions of things. So I was part of a group called, we created a group underneath the LDBC, the Linked Data Benchmark Council, to propose a language for property graphs. And we called that G-Core. And that's basically, we kind of just sat together and said, hey, if we could design a language from scratch, what would the ideal language look like? So we developed that. And then all of that went inside of the ISO, influenced the ISO standard. So right now, the same way ISO is the standard for SQL, ISO has now created a standard for the query languages and for property graph called GQL. And that's coming out soon. This is the first time that it's ever standardized another query language ever. Wow. 40 years. So that's an indication of how the world is going. Yeah. Graphs. So I guess for the learners out there, they've probably been exposed to the relational database or exposed to variations of different types of database, but mostly things that store data in rows and columns, right? Or maybe semi-structured data, perhaps. But the graph feels like it's a bit of a departure from these types of databases and storing data and querying it. Why would somebody want to use a graph database versus, say, a relational database? Yeah. So great question. I'd love to dive into this. So if you think about a relational database, one of the things is that you have to define the schema up front, right? So now, again, if you're starting to store JSON blobs in there, then you're really not really using the relational schema somehow. You're just kind of... So the point is, where is your application, where is the logic going to be embedded, right? So your logic of what things mean should be inside a relational schema. And if you have fixed requirements, you have a fixed application that you're building, then you use a relational database for that. That's exactly what it's meant for. You have the requirements. You know what it is. You know what are the types of questions people are asking. You can optimize for those query workloads and so forth. That's what a relational database is for. Now, when it comes to graphs, I think, first of all, there's the graph model itself, these triples, these nodes and edges, by nature, they're very flexible, right? It's something I can just go... I can easily add nodes and edges without that. Think about it in a relational database. I have to change the schema, update the schema, I have to go ahead and call them at the table, right? If something goes one-to-one, I'm going to make it one-to-many. I got to go refactor all these things, right? So there's a lot of work they have to do if there's a lot of changes that are going to happen, or if you have a lot of... You have some known requirements, but you know that things are going to be changing a lot. So the graph is a very flexible nature, just flexible by nature. So if you have requirements where things are changing, then already graphs is one of the reasons why you want to go and deal with things. The other aspect is that it's a minimal, it's a higher level abstraction. So they always think about computer science is always about finding different levels of abstraction. And as a technologist in computers, you kind of find where you're comfortable in your level of abstraction. And that's kind of where you dedicate your career at, basically. Or sometimes you become a compiler, right? I work great between these two levels of abstractions, or I'm an expert in moving between levels of abstraction. So you work in UI, right? You work all the way down at the bits and the chip level, right? Oh, this is how this all changes. Now, when it comes to graphs, I think it's just a higher level abstraction. So when you talk to users and, hey, how do you think about something? Let's go draw it on the whiteboard. You end up drawing bubbles and lines. So I think you have this higher level abstraction that you want to be able to kind of, what I call, bridge the data meaning gap. So I think if you want to be able to focus a lot more on what does this stuff mean, and this mean is changing, that's where graphs come in. And then also, if you have all this, let's call it just graphy type of questions, right? I want to know the path between something. I know if they're one of the cliques, I want to know the centrality around things, right? These are applications that are, I mean, graphy type of applications, recommendations, right? You want to go do fraud detection, network analysis around that. Then the other thing is that graphs is this minimal common denominator. I can turn any other model that I want into a graph. You can take tables, turn it into a graph. You can take JSON, XML, tree-based structures, turn it into a graph. When you do some NLP and you want to codify it, you're going to extract entities, relationships, turn it into a graph. So a graph is this common denominator that if you want to start integrating data coming from so many different sources, known sources and unknown sources, it just has all the great features around that. And then I would argue if you were trying to go do that inside of a relational database, you're going to be doing so many changes that you're going to end up building a graph model inside of a relational database. And therefore, when you start seeing that, you're like, well, you should have started with the graph. So I think that's why if you're doing a lot of data integration, metadata management, cataloging, that's why data in our world is on a knowledge graph because it's all about bringing metadata from so many different places. And you really want to manage all types of data. That graph is the ideal structure. So my long-winded answer to say, if you have very fixed requirements, then you create your application on a relational database. If you're bringing in data from so many places and things are going to change and evolve, then you're naturally going to be thinking about a graph. And one word that you use a lot when we hang out and one word you post online a lot is the word knowledge. What's a knowledge graph and how does it differ from a graph? So I think, so a graph, consider it just as the data, right? The points around how they're written and they're related. The knowledge really is capturing the meaning. Give me the first way to think about it. You're having a schema, right? And this schema can start evolving to have more richer information. So for example, you can say there are orders and there are customers, right? Okay, orders are placed by customers. Then I can start adding some type of cardinality. You can say, hey, an order can only purchase exactly by one customer, but a customer can purchase many orders, can place many orders. And then you can start thinking about adding more information around taxonomies. So for example, we have geographies and I want to know things about cities and states and regions and so forth. I have orders include the products that I'm selling. There's a product taxonomy around these things, right? So I want to be able to start connecting all that stuff together. I want to be able to then do some reasoning, right? Oh, if I'm interested in some, I don't know, in cosmetics, for example, right? And cosmetics is a super class. Lipstick is a subclass of cosmetics and stuff like that, right? So you start to get all those relationships, right? Those hierarchies in there and you can get much more richer and richer information. Now, and that's what I mean by knowledge right there. So if you think about just the graph per se, you're thinking about just like the raw data, but then the knowledge parties was actually bringing in all those semantics, that meaning, and it's connected to what users, how the users think about and are thinking about that domain. Interesting. And I guess we're filming this in 2024 and large language models are incredibly popular. And I know that you've done some research recently on the impact of knowledge graphs on large language models. You kind of walk the learners through the impacts of knowledge graphs on LLMs. So what's happening right now is that we're all obviously in the hype of LLMs. And I think what's really important to understand is that these LLMs are trained and they know, they quote unquote know, they do seem to know here, things that they've been trained on. And a lot of it is just the data and stuff that's on the web. Now, it doesn't necessarily know what's inside of your organization. Now, a lot of people are saying, oh, I should just then train this thing with me. I'm like, I should just train it with my data. I'm like, only a handful of people in the world can probably start training things and really you don't even, people don't even have the time or the money, the people to go do this stuff. So that's kind of really not that feasible. Now, even if you were to train that, your data is getting updated all the time. So you can't, so it's something not very feasible at this moment. But so what these LLMs don't have is the context, the semantics, the meaning of your organization. So a lot of the, what people are really interested right now with these large language models is to go chat with their data. And I think a lot of the chatting you've been doing over documents, over text, and being able to do embeddings and put those into vector databases, vector database, other database, database to manage vectors, right? So, but what happens with your data that is in your structured formats, the data that actually has rich meaning around that. So people, that's where the whole next to SQL thing comes around. And then a lot of the thoughts people are having is like, oh, I can just use LLMs to generate, give it a question and generate the SQL query for me. I think that's where people have been at, kind of their mind was at probably a year ago. And many of us were like, yeah, technically it could work, but how well does that work and how do we improve that accuracy? And the hypothesis was always, well, these knowledge graphs actually have all that context. So a lot of the experiments that we did late last year was to be able to test this question answering over using LLMs and saying, hey, if you wanna go answer questions over your SQL database, and if you add a knowledge graph, if you don't add a knowledge graph, what's the difference? Well, long story short, the research that we did was like three times more accurate if you actually are putting the questions in terms of the knowledge graph. And specifically, I mean, it's you're taking the question and you say translate it to a SPARQL graph query given the ontology, and you compare that, take that question and translate it to a SQL query given your SQL schema. Just with that very simple prompt, it was already three times more accurate using the knowledge graph. Now, this is just a baseline. This is just a starting point. And what we're seeing is so many people have been validating that work and showing, oh, they're not only just validating, they're finding ways to go improve it. Now, one of the things that I hypothesize why this is actually getting better, because we really don't know what's happening inside an LLM, is that language, right? LLMs are language. And we started off saying, hey, these are subject, verb, object. These triples are subject, predicate, object. So they have this kind of language in there. So that's one thing. And second is the relationships in the graph are first-class citizens, right? Something that you're managing, that you understand what this relationship is, and you're giving it a name, giving it semantics, meaning, while in a relational database, it's really implicit inside the foreign keys. And that's even assuming you tell it what the foreign keys are. And maybe the names of the codes that are being used, the column names are not really that descriptive. So you're like, well, I just have more context inside of the knowledge graph. And I effectively tell people that if you're trying to go add all that context to your relational database, you're going to be building a knowledge graph anyways. So invest in your knowledge graphs from the beginning. And that really is that context. And that's the engine that's going to be powering a lot of these AI apps that want to go talk to your structured data. It seems to be the case. I guess, closing out, is there anything else you would tell somebody who is new to graphs or knowledge graphs? And maybe some advice you'd give them on learning more about them. Yeah, so first of all, this is a paradigm shift. And it's a social technical paradigm shift. And I think, one, if we always look at tables, and tables are always in our lives, right? Then you think everything can be solved with tables. And so I think the quote I use in my book is the limits of my language are the limits of my world. So it changes hard. So my suggestion is, let's be open. Let's go try things. Let's try to get out of our comfort zone and try to figure out, see things outside of our bubble. That's number one. Another thing is, this isn't just the latest, greatest thing that people are inventing right now. Graphs have centuries, right? And the history of knowledge graphs go back a long time. And I encourage people to just look at, I wrote a paper a couple years ago talking about the history of knowledge graphs. And I think it's just a great pointers of where you can go back. And a lot of the things that we're trying to go do today are just things that we have been doing in the past. So just remind ourselves that we want to stand on the shoulders of giants. And let's not spend too much time reinventing the wheel and understand what's been out there. And it doesn't mean that what's out there is perfect. We can use it, but we should just go build on top of that instead of just thinking that we're just, it doesn't exist. And we start from scratch and be proud of building something from scratch when you've really just wasted time. Awesome. That's my honest suggestion. That's awesome. Well, thanks for your time, Juan. I'm sure the learners will get a lot out of this discussion. So thank you very much. All right, thank you for having me. Anytime.