In an era where data is paramount and AI is rapidly reshaping industries, organizations face the dual challenge of securing highly sensitive data while simultaneously fostering an environment of open sharing and rapid innovation. This session, featuring insights from Databricks Solutions Architect Luke Bilbro and a panel including data leaders from the World Bank and Petrobrass, unveiled a blueprint for achieving unified governance and enterprise sharing for data and AI.
“If you just decentralize without governance, it will become a mess.”
- Luke Bilbro, Lead Solutions Architect, Databricks
Struggling with data silos and slow insights? Discover how unified governance and open sharing are revolutionizing data and AI. Learn from real-world successes and get the blueprint to future-proof your data architecture.
Uh hello everyone. Welcome to the uh first public sector breakouts uh session. Uh hopefully you're all in the right place. If not, I encourage you to stay anyway because we're going to talk about some good content. Um, I'm going to be discussing uh unified governance and enterprise sharing today and I'll be doing about half of the time and then we're going to invite some panelists up for uh a discussion after I I conclude. Uh, first I do need to do the obligatory uh warning that some things I'll be talking about today are going to be a little forward-looking. I thought there was going to be some stuff announced this morning. Turns out it'll be announced tomorrow. And so um just know that maybe not everything is announced quite yet. Uh just bear with me on that. I am going to try to stick to what we have in the here and now though quite a bit. As a quick introduction, uh my name is Luke Bilbro. I am a solutions architect at Data Bricks and I've been here going on uh almost six years now. Uh and over that time I've had the absolute
privilege to work with essentially every team in public sector. I feel like uh a little bit off and on throughout the years, but I I've always been aligned from the start with the United States Postal Service. It's actually why I was originally hired uh at Data Bricks back in early 2020. And I think for right or wrong, I think the the connection I have with Postal Service and the work that they've been doing um since they started using data bricks is actually the reason why I was invited to uh give this session today. Um at least that's how I'm going to interpret it. So there's going to be a little bit of a USPS lean to the to the conversation to provide some context. But before I get uh too far ahead of myself, why don't we do a motivation slide? Um why the focus on governance and sharing as a unit uh in this session? Um some of that is going to be pretty obvious, especially the governance side, right? I mean governments and research institutions are going to produce a ton of data. uh it's very very high sensitivity and so of course keeping it all secure is absolutely paramount but uh and that and I would say more so than in other industries that's true but uh on the other hand very much like other industries uh the government organizations are actually trying to do similar things they're trying to modernize oops wrong button they are trying to modernize systems cut costs leverage the latest tech and uh deliver results faster and here is where the opportunity comes in for sharing. And I don't just mean sharing data. I'm talking about sharing insights, successes, failures, advances, everything. Because if departments within organizations were more free to share assets that worked, all departments would be able to advance uh together and faster. But the trick is how do you do that in concert while factoring in bullet one? So that's another reason why we're doing these together. And then I guess you know maybe there was like this you know relevant executive order recently that probably was pushing for increased data sharing and I would wager that that's probably has something to do with the fact that we're talking about these together today as well. But what are government organizations to do now? They're in a tight spot. They need to they need to uh modernize fast. They need to leverage Genai. Uh but they need to keep um security and governance in in um mind here. And that's where I want to brag on the postal service a little bit and all the work they've put in to date and all the work that they've uh continue to put in. And I think their story is one that might uh be interesting for everyone here. Uh it's definitely topical. Um certainly on on point here. So I'm I'm going to start with a a very brief uh summary and a bit
of a timeline of postal's journey to a sharing centric enterprise architecture which is what they are rapidly standardizing on presently. Um they're calling it the unified analytics platform which or USPS UAP for short. Uh of course you know PubSack loves their acronyms. Um but that's not how all things started. Uh in 2020 we actually kicked things off in a completely different method. Uh the whole idea back in 2020 was a heavy focus on hardcore data engineering and data science. The goal at the time was to lift and shift and then optimize some existing Spark workloads to improve processing performance and efficiencies and cut down on the dreaded maintenance overheads, which is honestly pretty standard fair for data bricks back in 2020. Uh by the end of 2021, USPS had actually started expanding the vision that they had and testing a newly released product data brick SQL. Uh the idea there was they were just trying to see could it optimize some existing BI workloads that they had been running for a long time. And I I will always remember one one specific use case. There was a BI report that I think ran on Oracle and it took long enough to finish a couple weeks I think that by the time it did finish they had to immediately start it again in order to get the reporting SLA uh in time. And so there was almost no window of time to uh improve the uh the performance or or iterate on the report. And uh I remember they eventually after moving to data brick SQL that time dropped to just a couple of hours which just completely revolutionized the way the team uh thought about um optimizing and working with their reports. So that momentum continued to build with kind of each smaller success. Um in 2020, Unity catalog GAD and uh postal service had heard about it in 2021. They were very eager to get into it, but by the end of 2021, the the planning was in place for how to migrate to Unity catalog. uh and they was it was also around this point that a vision had started percolating around the enterprise um particularly in the CDAO CDAO organization and the idea being asked and kind of moved around was could Unity catalog be a central hub for enterprise data assets uh at USPS. It was still pretty early and so there was no answer yet but but the ideas had started as far back as 2022. Uh by 2023, proper plans for the UAP were in place. The UC migration was in full effect. Uh this was going to become the trusted location, the source of truth for USPS enterprise assets. Um and it was also around this time where we hit one of our first major tests and successes thankfully uh of a data sharing centric ecosystem. Uh it was in
mid 2023 when a new type of shipping fraud had been identified and prioritized for solutioning and some kind of response. Um and very very quickly the the postal team identified that they needed more data to properly scope the problem. Uh but when they figured out what data they needed, they realized it was actually two main data sets that resided in two different uh data stores on prim and that was going to be a problem if uh they didn't have a fast way to get the data out and uh combine it. Thankfully, both sources were already being ingested uh into Unity Catalog by two independent teams working in two different workspaces with uh no overlap. But since the data was already there in Unity, it was already implicitly sharable. And so by the next morning, a SWAT team of data scientists had been deployed to one of the workspaces and had full access to all the data they needed. Uh and so that absolutely moved the needle forward and kept the the uh progress going and the journey has continued from there. But I don't want to give away what I'll be talking about in the rest of the talk. So, um, why don't I moving move forward
and say knowing what we know now with our perfect hindsight, what is the blueprint to mimic USPS? And I think that's going to be the goal of the rest of the talk here. Um, and it boils down in my mind to three learned lessons um that we are continually evolving on at Postal Service. First is that you have to be able to prove that you can control your data assets. And there's two parts to this. It's great to have a tech that can do it. Uh, but you have to have a team that's trained on it and they have to have an opinionated plan. Uh, in other words, you need the right tool. That's definitely part of it. But you also need a dedicated and enabled data administration team. And depending on how your organization is organized, that team may be completely different from your platform administration team because we're not really talking about cloud infrastructure here. We're talking about data organization, security, and ownership. The second point is that while it might be tempting to immediately start sharing data out as soon as it's available, you're probably better off spending a little bit of time building trusted high-quality assets first, ensuring that those are maximally discoverable. Now, to be clear, everything is on the table for sharing when you're using Unity, but it's probably not where you should start. We would we would recommend you start small with quality assets, and that'll help build trust with the collaborators and the consumers downstream. And then finally, my last bullet point, uh, which could go without saying, but I'm going to say it anyway, is that you need to standardize on an open sharing framework. You should be visionary here. You should try to imagine all the troubles that you might hit in the future, all the future scenarios, because sharing will definitely unlock short-term gains as teams are able to collaborate faster and more closely. But a proper open sharing ecosystem will also futureproof your architecture. when the next big tech thing hits, you'll be ready to move with it because your data assets will be open, organized, accessible, governed, and most of all, yours to do with as you will. So, I'll jump into each of these three bullet points a little bit. I'm going to quickly scan through governance. I think this is maybe the one everyone knows best. You've heard a lot about it. In short, data governance really just boils
down to one key aspect, establishing trust uh in data because only if consumers trust the data will they actually care about or buy into using it. And that's always going to start with ownership. If you have to start somewhere, start at the top of this cool looking spiralally funnly thing. Um there has to be someone accountable for data, someone who makes the ultimate decisions. And after you have that accountability in place, then you can start layering on all the other goodies that you would need. uh compliance uh measures and quality checks and eventually facilitating the usage of the data. But I'd like to think about this a little different. We we can break governance down into maybe two different parts. And in one way they they kind of behave like uh different branches of the government. Think of it like legislative and executive branches here. Um pure governance is really about authoring the rules and the compliance uh definitions, the the framework. um it's more like the legislative uh side of the government whereas management is actually more about enforcing those policies. It's more like the executive branch. Um so governance is the framework and management is the implementation and of course they work together in harmony hopefully. Um but uh maybe not always. Um and you know I think the judicial branch probably fits into this analogy somewhere. Um but I will leave that as an exercise to the audience for later. Didn't have enough time to come up with that one. Uh so just like the two branches of government apply to work together uh to facilitate government in the real world, you need them to work in concert uh with your data. And I'm going to come back to this analogy probably too many times. Next, I want to uh map out a standard
journey from start to finish. And when I say standard, I mean like what we normally see out in the field from our customers. And when I'm talking about a journey, I'm talking about going from relatively no governance position or a naive position all the way through to a pretty mature way of thinking about governance. Um, almost always the way that it tends to happen is in a fairly bottom up fashion. Uh, and as a little hint, if you want to try to maybe shortcut your way to uh, a more advanced maturity level, you should at least try to think about things from top down, even if you're going to implement them from bottom up. But more about that in a second. Let's go back to what we normally see. We we would normally start here at the bottom uh, under this blue large data management ring, maybe below it actually. And you'll have a team that starts using data bricks. They're ingesting data. They're making great progress. and other people catch on, other teams catch on, they spin up their own data bricks and they start doing the same thing. More teams do it and you start scaling out horizontally. Uh but at some point all of these assets need to be uh mapped together. You really need to organize the chaos at some point. And thankfully with UC you get that out of the box. All that technical metadata is automatically captured as soon as you're using UC. This was not the case before UC with Hive Metas. And for those of you who use the Hive Meta store, uh, good on you and welcome to Unity. Much better. Um, it's easy actually to stop there though, uh, because you're getting all of this out of the box. But I can't emphasize enough, don't stop there. Uh, in fact, it's a trap. Don't stop there. You will soon end up with thousands of assets across multiple cataloges, all technical. And while everything that is in there is somewhere that you could find it, it's findable. It's really hard to argue that it's discoverable. Um teams that are working on data strategy or maybe cross functional teams when they're trying to look at other people's data, they're really going to struggle to understand what the data is. They're not the team that produced the data. They don't have the fundamental knowledge to figure it out. They're going to try to figure out, do I need these tables? What do these other tables mean? How do I find what is really relevant to me? So the next step on the journey is actually I think the most important one. Uh it is to push your
thinking towards data products. And data products can mean a lot of things but in general it's just an abstraction layer that aims to help solve some of the problems that simply technical cataloging is going to uh the shortcomings of just strictly technical cataloging. Um, and maybe the simplest way to do this, uh, of course there's a lot of ways, is to start grouping assets together into logically related sets. And once you do that, it be it becomes easier to make them discoverable as a group uh, and also to make them accessible or inaccessible as a group. The last step on the journey is to align the data products to strategic goals. And that's normally done by mapping the technical lingo that they're usually written in to some standardized enterprise business language. That's really where the concept of the semantic layer fits in. But uh although this is the journey again I hinted at it. I I really want to focus on the data products because I think if I had a key takeaway for you all that are embarking on this journey, I think the hardest step is to get out of that uh purgatory at the bottom. Don't don't rely on just what comes out of the box. really start thinking about things in terms of groups groups of related data and what are things that are easily consumable. So maybe just a bit more on that. This is really at the data product level. This is really where the importance of ownership uh sets in. The data products maybe potentially unlike the raw data that you have in your technical cataloges have to have a clear owner and they probably are going to need a lot of metadata to help people understand what they are. You're not always going to be there as an owner to explain what these data sets are. That metadata might be done through descriptions and tags, maybe table tags or column tags, but the information might explain what the refresh schedule is or whether the schema is fixed or dynamic, um whether the or what data lineage uh it comes from, what are the actual source systems of record, um what's the data sensitivity level, and um really a ton more. It really depends on your legislative department that you bring to the table um of of what you're going to put in there. And maybe to put things in perspective, while while you might start with say tens of thousands at some point of technical assets, you should be thinking an order of magnitude fewer when it comes to data products in general because these are things that are derived from your technical cataloges and um are more like the tip of the iceberg. These are going to be maybe gold or post gold assets if you're thinking about the medallion paradigm. Moving forward, I get to finally maybe
talk a bit more specifically about Unity Catalog. And back to my um questionable analogy, Unity Catalog, I think, fits the role of that executive branch of those responsibilities really well. It comes with tons of data management features um and compliance features, notably auto collecting of lineage and audit logs as well as some okay well this is this is the part where I have to be careful uh some things I thought were going to be announced this morning um but I said it was forward-looking so as well as some newer Aback uh and data classification features that you may have heard of in today's or tomorrow's keynote um and even our first fora into semantic layers with um metric views so um be ready for that tomorrow But uh whereas data bricks and unity catalog are providing you with the strong set of tools that you need the legislative branch is still needed and that will come from your teams. That's the only place it can come from specifically your data strategy teams that keep focused on this kind of data product level of thinking. UC is ultimately just the executive branch the enforcement arm but without a legislative arm to with opinionated views on what data governance means for your organization you're not going to get much farther than that technical data catalog that I showed earlier.
So, uh, I'm going to pivot now to my third and final bullet point of the, uh, USPS blueprint, and that was for sharing. Uh, and I'm sure you've all heard about delta sharing. It's been around for a while, and honestly, it's still a fantastic option when it comes time to actually share out your assets that you've built. Um, the open delta sharing protocol can be used with any compatible client. Consumers don't need to be data bicks users or licensed with data bicks. And whenever consumers use Delta shared data, it uh it's automatically there's no compute cost in the provider. It's automatically logged. There's lots of security features. So all the same goodies that you had before except now you can delta share more than just data. You can delta share uh tables, views, materialized views, volumes, notebooks, models, all of that is now sharable with Delta Share. So definitely a great way to uh share throughout your ecosystem. But while most of you probably knew about Delta Sharing, I'm way I'm willing to bet a lot of you didn't know about our REST APIs. There are now uh open APIs directly into Unity that allow yourselves and your collaborators to really choose any tool that they want uh and interact directly with Unity. And when I say interact, I'm talking about reads and writes. And they can be Delta clients or they could be iceberg clients. Doesn't matter. We support it all. Uh, and so this is this has been something I've been waiting on for a long time. Uh, and I was really really happy to hear that we have these out now. So, multiple ways to share, multiple ways to be open, multiple ways for everyone to choose the tool of choice. We'd love for you to choose data bricks, but we have to earn that right just like every other tool out there. So, uh, just to overwhelm your eyes, uh, lots and lots of, uh, tools out there that that can be used with all the different sharing protocols. But um to start summarizing now before we hand it over um I talked about USPS's journey to the same state. I talked about the importance of data governance uh and the idea of governance versus management. The idea of focusing on data products um and also maybe eventually to data domains um and then finally just now about sharing. But the final step is to
maybe justify the title of this session which is on unified governance. And probably some of you have already kind of seen a little bit of this, but uh Unity Catalog from the start has been a great tool for governing the assets in your lakehouse. But what about the data not in your lakehouse? What about the systems of record that for a variety of reasons are never going to be ingested? And those reasons are numerous. Some data really needs to be external, plainly put. Well, with Lakehouse Federation, which has been improving and improving over the last six to 12 months with the number of sources, you no longer have to worry about that. You can fully uh unify your governance just in UC. Leave the data where it is, register it as a foreign table, and then govern those assets just like they uh just just like they were colllocated with all of your lakehouse assets. Um, Federation already supports a lot of sources, a lot of the big boys, Oracle, Teroda, Snowflake, and more. Um and so this is how you can extend your boundary. And that brings me to my basically my last real slide. I have one joke slide but one real slide uh where we can summarize this session with kind of this uh beautiful image saying whe wherever your data products originate whether they're ingested or federated um you can actually govern them all in one place and whether your collaborators use data bricks or any other tool they all come to one place to access that data. So this is unified governance that is built for enterprise sharing. And uh on one last funny note, I need to thank our design team for making these wonderful slides because about a year or so ago, I tried to convey almost this exact same message to Postal and the best I could come up with was this. Um and so sorry about that y'all. Um and thanks Data Bricks team. And uh in case you like the vertical flow instead of the horizontal flow, there's another one if you wanted to grab a picture. But, uh, with that, I'm going to close it out and, uh, and
bring up my my colleague, Molly, uh, to kick off the, um, the chat, the fireside chat. Take it away, Molly. Thanks, everybody. Thanks, Luke. Got it. Thanks, Luke. That was a great great segment. So, today we're going to follow up uh, Luke's great segment on the technical content, and we're going to bring up two great panelists. We're going to hear about the World Bank and Petroboss, how they did pretty large-scale deployments of data bricks and how they enabled governance and enterprise data sharing. So, please join me in welcoming Siraj Cotti who's the data leader from World Bank and Marcelo Dotto from Petro. Come on up, guys. Hi. So, Sesh and Marcelo, to start off the conversation, could you share with the audience kind of a little bit more about your role, your organization's mission, and the key challenges that you each are trying to solve? Thank you. Am I Hello. One minute. Hello. Oh, there we go. Yeah. Thank you guys. Yeah. Thank you databicks. Thank you Molly uh for having us over for this session. So my name is Suresh Kaudi. I work for the World Bank. Um just so that you guys know the mission of the World Bank is to elevate poverty and improve shared prosperity for the bottom 40% of the people living across the globe. So what we do as data, what we do to support this mission is to enable our data, enable decision making, enabling providing deep insights into the data. So I have been working in data all my life for two and a half decades. Um in the initial stages I was one of the data junkie and then later on u moved on to become um formulating the data strategy back in 2012. We formulated a strategy back then based on the technology set. Um we had one monolithic system which had pretty much done the AI BI loads very well but we had pockets of um other areas that we relied on several tools for data virtualization and so on so forth. Currently as part of the evolution of the cloud we have formulated formulating a data strategy which is more interoperable uh where we have different best of the breed tools coming together to solve the problem of making sure that the data is accessible to eliminate the silos unity catalog that is being talked about largely we leverage that we use the data sharing mechanism which is the virtualization standing up the strategy for a IML workloads And that apart from that um I also have participated in uh developing a strategy for the Caribbean islands where they had to formulate the payment system data was in silos. We had to formulate a data strategy for them where they could collect the taxes from the citizens and also worked with the country of of Georgia to formulate their their data strategy as well uh to eliminate the silos the fragmented data sets across and bring um everything together. Thank you very much. That's incredible. Hello everyone. Thank you Molly for the invitation. My name is Marceloto. I work at Petrobrass. I don't know if everyone knows Petro. It's it's a Brazilian oil company. H Petro is the biggest company in Brazil. It's one of the biggest companies in energy and oil all around the world. Uh we work from exploration, production, refining and commercialization of oil derivates. Um we have a a a leadership in deep exploration and ultra deep exploration. It most of it because of our 50 years expertise on that studying exploration. uh and at the moment our main goal is at the company is like uh increasing operation efficiently because to reduce costs and also to reduce carbon emissions this is very important the company at the moment I'm at Petro Brass for 30 years 13 13 years 13 years at Petro Brass I was thinking I was like he is much older than I thought not You look amazing. Whatever water you have, I will have that. I met a Petro Brass for 13 years. I work with data for almost 15 years. I started on Petro working with data. I I was on infrastructure team for data. Then I started to develop uh for data too and at the moment I am leading the data integration platforms team. We have uh a lot of thing to talk about data. on the next two questions I can. Yeah, absolutely. Um, so that's a great leadin. So, prior to the data bricks deployments, can you can you each talk about what the key data and AI challenges that you were facing and following the data bricks implementation, what was enabled as a result of that? Okay. So, I'll go first
and then my colleague will join me. So before the implementation of data bricks as I said we formulated a strategy in in 2012 with a near-term and a long-term strategy to to have solved the the data silo problem. But what we had was a monolithic system worked extremely well against um the BI workloads. What was missing prior to data bricks in that was we didn't have a unified layer where we could get all the data together. we didn't have or we didn't even think about uh unstructured content which is you know the documents the files the videos the images and and so on so forth and we didn't think about metadata as one of the critical aspect too with these loads we we didn't think about AI these were all the missing segments in the previous strategy with the emergence of cloud it's very clear that we had to bring both the structured content and the unstructured content together a house framework is the answer to that where we could bring in the the structured content and the unstructured content together and have a unified governance layer on top of it. The second challenge um the second thing that happened after the data bricks um implementation was the unified governance piece the unity catalog that is talked about everywhere where we could bring in both the data and the metadata together providing them the options of discoverability data quality cataloging which are key essential pillars of our data strategy strategy that were missing previously and we had to bring in as as part of this. The third aspect is AIdriven right everything is about AI this space the boardroom discussion the management meetings the the meetings at the ground level is about AI so that with data bricks we were able to bring AI to the front we were able to stand up um different use cases the NLP NLQ the chatbot options the narrow AI options compound AI use cases the agentic bots that are being talked out. Finally, lot last but not the least, we were able to also do lot of MLOps on our data to get immediate answers to lot of the questions that would take for us maybe a year or so to build using a planning application. We were able to quickly ingest our data and and get the answers to us. Finally, last but not the least, it's also about real time having different data, right? We were only focused on structured data. The streaming analytics data that is coming in the different niche use cases that are coming uh also plays a key role. Also the optimization of CL uh costs we were able to significantly reduce by having only um pay as you go option as opposed to keeping the 247 server options. And again there are a lot of options that we when you move to cloud infrastructure as a service we don't have to worry about you know managing patches, upgrades and so on. and so forth. It's done automatically for us and we can truly focus on delivering value to the clients and that is one of the uh greatest um benefit that we reap out of it. Thank you. Absolutely. Thank you. Well, uh I believe it in Petro Braz it was not very different from World Bank but um on the past we have some data silos on Petros and we had a problem of the difficult to to to deliver to the business areas data products uh sometimes we had like months to to deliver a data product. It it was uh very difficult to do that. So like four years ago we we decided to to to to move to a data mesh strategy. Uh we created a data platform. We designed one and it has a lot of components but the core component of the platform is data bricks. We have components another components for ingestion. We developed uh a marketplace or a portal for the end users browsing, navigate and request access for the data and all of the workflow for giving permission for the user is done on the portal. But the core component is data bricks. Uh I believe that this strategy uh gave us more speedy because uh we decentralized the development. We we have a lot of teams working on data nowadays and also we empower the business units to to start developing independent from from IT and it become a center of excellence on helping defining processes governance uh and implementing the the features that the the business areas need to to create their products. That's terrific. Um you mentioned empowering kind of the business owners. As you do that, what were kind of the key business values or metrics that you saw as a result of empowering the data end users and democratizing data at your organizations? That's a great question.
But let me take that answer and my colleagues. So first thing is um there are several if you look at your strategy of how you want to enable the business and provide value to them. I would categorize into four different sections. The first is the BI workloads which you run today. This the standard reports the structure of running dashboards uses. The second part is the AI workloads which is very critical right now because AI is the one that is that is an enabler right now. How do we augment our data? How do we augment our metadata and provide the AI workload um for that? The third aspect is the self-service that is not going to go away but I think with the invent of chat bots the genie rooms and so on so forth coming with data bricks announcement today hopefully that that piece will uh will pretty much shorten and the last but not the piece is the data data products we talk about our application delivery teams that rely on data data product is the answer to most of the questions and today the way we empower the business and made them look forward is to come through a data product which is already cataloged It is searchable. It is navigable. You can drill through provide the lineage aspect of it. You can put in the scores. Uh you can have it show the lineage of how where things came from. And finally encompassing it as a complete data product so that the users don't have to worry about anything but just take that consume that layer and use it. I think those are some of the the great inventions that are happening and again it's data bricks all over for this whether it be unity catalog whether it be data serving layer and so on. Um again it's not very different but I I can I can list here three aspects that I think that are very important. Two I have already said that it's faster time to value. Yeah when we decentralize and empower the teams they can deliver it faster. Uh and when we empower them, they they they they feel free to to create more data products on the on the on the on the environment. And the third and very important in my opinion is that we need to bring governance in all of that. If you just decentralize without governance, it will become a mess. Yeah. So unit is very important in our implementation. We decided to to implement unit on Petro just when it was released. We migrated right right in the beginning and I I believe it was a a good decision for us and that en so with this unified architecture we could get a lot of value for business. Absolutely. Um, looking forward, how are you each thinking about the role of AI? So, generative AI, uh, Gentech AI, uh, in your organizations and are there any near-term use cases that you're most excited about? Again, a loaded one, but we'll we'll
try. The next gen is all about AI. As I said, the boardroom discussions, the management discussion, the CIO's discussion. If you're not in AI now is not the fear of missing out. It's not FOMO. It's about FOBO where your fear of absoluting out yourself. You will be gone out of business if you don't embark AI. So the next gen of AI that is coming it is the agentic approach the way the streamllet application works or the the multi- gen room options or the multimodel options. I think there was an announcement made today about how easy it is to stand up this multi- aents and how each of them talk together. So I see the future of AI um as agentic approaches coming together. There are three main aspects that I feel um are are very critical to implementing AI. One is your business strategy where they define what value they want to get out of AI. If that is not there, I think you are starting off on on a not on a level ground. You'll start doing something which you're not going to meet. And the the the second aspect is the value stream that you get out of it. Right? What value are you doing? What there there could be many use cases for some organizations but picking the value stream out of it and doing it makes it real valuable. The third aspect is the the preparatory phase of I call it technology readiness. In order to be AI first, you need to have a data strategy. Without your data strategy, your AI is going to be in a vacuum. There are going to be silos, defragmented data sets. You will push data in what is needed for that application and you you'll always be struggling. Also, you need to make sure your technology is efficient. So you can lift and scale your things not just for one or two use cases, but you need to be able to scale um with hundreds or if not thousands of use cases. You should have the ability to to push realtime data, ingest the AI workloads, be able to do narrow usea narrow AI use cases, the chatbot, NLP, NLQ options. Um AI as a feature is become very predominant in in everything what you deploy. When you deploy a data product, it's important to have data uh AI features embedded in it. And also do augmentive AI. For example, the metadata that you want to go, you approach the business users, you give them the fields and tell them to put in good luck with that. With all the priorities that are there, it's going to take forever for them to do. you need to leverage augmentive AI where it can prepopulate already the definitions for you and then the business users come in and correct. One of the most important thing also is to keep your architecture more interoperable. You should be able to do plug and play. With the emergence of um the chat GPT or cloudy or B there is going to be change every single day. your architecture needs to be nimble, flexible, scalable, robust so that you can plug and play and ingest each of the things um as needed. Like I can give a classic DBRx model that came up was the top one for two weeks and then something took over. So we don't want to hardwire our architecture for AI in a way that you're always struck or bound to what you need to do. You need to be open, flexible, interoperable, more importantly to have your AI strategy right and done. Absolutely. I couldn't agree more. Nice. First of all, I agree with everything. Yeah. I'm I I'm not exactly from the AI team. I'm from I'm a data guy and um I I will I will focus on the aspect of data engineering. uh I believe
that uh it's it's proved that AI will help will help on business but for data engineering we see that the AI is helping on at the moment already with some repetitive tasks we can automate that and and deliver uh pipeline is much more faster than before. uh we we on the short near term we are planning to use AI for data observability and data quality to help better data because we need a uh good data for for having a a good result on AI and also we we want to scale insights for our our data engineer teams so we could we could uh accelerate more more the development of data products. Absolutely. Well, those are two great perspectives. Um, I want to thank you both for being here and thank you to the audience for for joining us too. Um, if you'll if you'll join me in in thanking our panelists.
Avoid this common pitfall
End your report nightmares
No data left behind
Trust is paramount
Secure sharing, faster results
Automate data management
Prove your data control














