In an era demanding agility and efficiency, government agencies face an uphill battle against outdated data infrastructure and fragmented information silos. This session, 'Reinvent Government in a Data Intelligence Era,' presented by Eric Papawitch of Databricks and Ricky Aurora of Minnesota Metro Council, unveiled a transformative approach using the Databricks Data Intelligence Platform to overcome these challenges and accelerate the adoption of data and AI in the public sector.
“We are only able to get to 20% of our data. Our data is we can't even process all of that data that's currently coming in. There's no capacity, there's no horsepower processing power at all to do any of this stuff.”
- Eric Popowich, Senior Solutions Architect, Databricks
Government agencies are drowning in data silos, hindering efficiency and mission delivery. Discover how a modern data intelligence platform can unify fragmented data, accelerate AI adoption, and revolutionize citizen services.
Good afternoon. Um I think it's time for us to get started. So um welcome to this afternoon's session, reinvent government in a data intelligence era. Uh this is split. I'm going to do approximately a 20 minute presentation and then I'm going to pass it over to um awesome to lead a uh fireside chat. Get the clicker here. And I do have a couple of uh forward-looking statements before I start. There are some features in the presentation that are not available in every region in which data bricks is deployed as a government presentation. I see some smiles here, but um just wanted to point out that there may be some roadmap features discussed in the presentation. My name is uh Eric Papawitch. I'm a senior solutions architect. I've been with data bricks for approximately four years. I primarily support uh customers in the uh federal civilian territory. Prior to this, I was in a similar role at Oracle for six years. Uh supporting Department of Defense. I've worked at Splunk in a similar role, supporting Department of Defense for six years, and if anyone can remember Vignette, I was there for nine years building a lot of the early federal uh portals. In the presentation today, um we're
ambitiously going to discuss how we can reinvent uh the way government does data in AI with the data bricks uh data intelligence platform which I think is pretty topical because for at least the last couple of administrations this has been a top government priority. uh with the IT modernization act uh there was a lot of emphasis on moving from on-prem applications uh to the cloud to improve scalability performance and have better cost models. Federal data strategy was an attempt to improve the federation of data and make it more accessible. More recently, the AI executive order has um made an attempt to accelerate the adoption of AI within um public sector. And then most recently within the last three months with the current administration, right? Uh very ambitious stopping waste, fraud, and abuse uh by eliminating information silos executive order. So, we're going to try and touch on parts of all four of these in this uh presentation. And I think we all agree that uh
government struggles with both uh data as well as AI. If we look at all of the government services listed on the left side of the screen, uh they they typically have the same architecture. There's a transactional system which is responsible for um receiving service requests or benefit requests or payments. uh there's an adjudication process to uh deliver that service or mission and then um that data is typically moved to an entirely different analytics environment. There may be some historical analytics maybe compliance analytics fraud analytics run on the data and then that data is transition potentially to a lake where there may be in some agencies AI performed to um uh do more predictive analytics. uh regardless when we look at these um individual application silos typically at a minimum the data estate is fragmented. So the transactional environment separate from the analytical environment and then in the case where there's a lake typically there's multiple copies of data that are transitioned to the lake and then we look at the agency as a whole um it can usually work look much worse because you can have every application having its own data estate making it very difficult to perform agencywide analytics very difficult to um understand the security posture of the entire environment and um resulting in many different uh copies of the data. So what we're going to propose
as I alluded to earlier was reinventing this architecture that supports both data and AI with the data intelligence platform. Uh there's two core components uh that we're going to discuss. The first is obviously the data lakehouse. So, this is going to be a cloud-based architecture that provides an open unified foundation for all of your data. And then talk about how we're integrating in generative AI to improve the efficiency in which data teams can access accurate data and build predictive analytics. Um, and then the two combined together are the new um data intelligence platform uh concept that data bicks has provided. So first kind of table stakes want to talk about unify and own the data. So with this architectural concept we're
proposing a lakehouse architecture. So this is uh something that data bricks developed in 2020 which is very popular uh currently. Um, it proposes using cloud object storage as the main repository for all of your analytical data. Um, and doing so in the agency or government cloud account. So no egress of data out to a SAS provider or outside of the boundary which it controls. obviously improving the security posture of that data estate infrastructure but also um uh providing a much better cost model to some of the legacy um storage models we've seen in the past. We're also proposing uh the use of an open storage format whether that's Delta Lake or Iceberg. The two seem to be the market leaders right now. This too is highly strategic because it reduces um the vendor lockin that's very uh obvious and apparent right now in existing data silos. It also does not lock you into data bricks because there's an entire ecosystem of analytic tools that support both of these. And then um in terms of lakehouse
principles for modernization, there's three I really want to drill down on. So the first is that um as we build out the lakehouse and build our ingest pipelines, we want to ensure that we're curating the data to offer trusted data as products. So in this environment, uh we want to ensure that we take the raw data from our transactional systems, enrich it, uh clean it, make sure it's high quality, and then potentially normalize it to improve the performance of any downstream analytics. Um, the second principle I want to um, dive a little deeper in is adopt an organizationwide data governance strategy. Um, I'll talk about that a little bit later, but essentially once we have our clean data, our high quality data, we need a way to be able to um, enforce proper access controls. We need to be able to audit access to the data and we need to always understand the data lineage. And then the third principle which we'll talk about is not only using the uh open storage format but also using an open interface so that we can share both intra agency as well as potentially across a wider um um part of mission partners. Okay. Okay. So if we look at the
next core design principles, we want to make governance a strategic asset and this is necessary because for most public sector workflows, we're dealing with sensitive or highly confidential data. Okay. So there is the requirement not only to have strict um control within our cloud infrastructure but also access control perhaps even fine grained access control on the data itself. And then we need to be able to perform continuously or we need to continuously monitor that access so we can um ensure we have compliance. But we also need to democratize the data. There were at least um two of the earlier executive orders or acts that uh talked to open collaboration, data sharing, making sure that not only data teams within the agency have access or appropriate access to the data, but we're able to share out data with appropriate access to a wider uh set of partners. So, data bricks is solving this problem with Unity catalog. So this is the governance layer that we're applying directly on top of our lakehouse architecture and um is giving us the best of both worlds. And from a feature perspective, we're providing governance on structured data, but we're also providing governance on any unstructured data. Think of form data that may be part of the intake process for your mission delivery as well as governance directly on top of the AI functions or AI models that you may develop as a part of your uh lakehouse. Uh there's built-in auditing on all of this uh data as well as the ability to apply the fine grain access controls that we need. In addition, we can federate um out to external data sources if we need to bring those in for analytics. Realistically, especially in the early stages of a lakehouse implementation um at IOC, we're not going to have all the data we need to perform our analytics. So, we can federate out to those uh existing data silos to bring those in, govern them, and make them a part of our um data estate. And then we can also provide through delta sharing a means of uh giving secure access to the data that we have in the lakehouse. We're also uh continually improving the
product by integrating in uh artificial intelligence. Uh two things I'd like to point out. The first is the ability to uh classify our data and AI assets. Examples of this being we can add a description automatically to a table or comments to columns. And this is obviously important for data teams. If they're searching for data, they want to be able to find the data sets that they need. When they find them, they want to understand at a column level what's in the table. But also as we build additional AI agents, the AI agents need that metadata to understand how to use the data. Okay. So serves multiple functions there. So we can have this uh auto classification done with AI. So we don't have to have our data stewards necessarily go in and manually write it all themselves. They can just act as the um um the auditors of what's been provided. We also can autodetect PII. So as a pipeline executes, have Unity catalog uh through pattern recognition determine potentially if there's social security numbers or anything that may need to be obsucated or have stricter uh role-based access control applied to it. Okay. And then this leads us to the
ability to support multiple different versions of data mesh patterns uh essentially agency as a domain. So we're seeing um obviously at the agency level if the agency is the data domain the ability for them to uh integrate into a wider mesh using delta sharing. We can have sub agencies or programs within a department act as their own data domains data stewards and then share out data either uh intra agency or to other uh mission partners. What's interesting about um delta sharing is because it's an open protocol, we can do this within the same cloud region. We can do this across cloud regions within the same hyperscaler or we can do this across uh cloud providers. So AWS to Azure it's an open interface once we have the share defined if the appropriate privilege via token is granted that share can happen. What's also nice about this protocol is that there is a large ecosystem of tools that support it. So an external party that doesn't use data bricks can even be a part of the mesh as long as they're using a tool that supports the sharing protocol. This aligns nicely with the EO we talked about earlier. Stopping waste, fraud, and abuse because we can eliminate data silos to a certain extent making unclassified data accessible via this type of pattern. Okay. And then finally to uh drill down
a little bit on how we can take this govern data estate and accelerate datadriven citizen services. The first feature of the platform that I'd like to talk about is data bricks AI BI. two core components obviously dashboards um which give analysts the ability to visualize data sets as well as Genie right which is a conversational agent uh built to give data teams the ability to ask natural language questions get answers back okay so how does this work um from the dashboard perspective guess it's not rendering or let me go back bear with me for the middle here So from a a dashboard perspective, so AI is integrated in three ways that I'd like to point out. So the first is to build the data set. We can use data bricks assistant to use text to generate the SQL to build the data set that backs each visualization. At the dashboard level, we can similarly use an integrated agent to not only uh select the appropriate visualization but also then configure it in the dashboarding template. And then once the uh dashboard is published, we can introduce uh different predictive analytics models to uh improve an analyst understanding of the data. The example we're showing on the far right is the integration of a forecasting algorithm. So you can take that backing data set then run the forecast on the data to predict out what the ensuing time periods are going to be from a a result uh result perspective and then with Genie um three interesting things about this so this is an agent system it's comprised of a number of AI agents uh it works and it works really well because it can leverage context from Unity catalog right that platform metadata we were talking about earlier. So, not only are we telling it what tables to use as reference, but also it has all of that additional metadata so it can provide more accurate answers. And then when we build uh Genie space, we can put plain text instructions on the space to give it additional context how to calculate revenue, how to determine um uh certain uh answers that may not be intuitive based on just the tabular structure. So this is effectively a way to um fine-tune the way Genie uh operates on the data but in plain English. And then finally um as data teams are using Genie there is the ability to incorporate feedback. So there's simple yes no if you haven't used it before there's the simple you know is this answer accurate yes or no. So that feedback is incorporated but you can also give additional context if the answer does not look uh accurate. So there's an iterative uh improvement cycle that's improved inherent to the way the um agent system is implemented. So when we combine the two, we can really improve uh service delivery. This is a simple example, but um when you build dashboards, right, it's almost impossible to anticipate all the questions that are going to be asked of the data. You do your best to uh represent um the data in a way that's meaningful to the data teams, but there's obviously always additional questions. So with Genie integrated, we can shortcircuit the um uh feedback loop or development loop that's required to get those questions answered. The analyst can simply ask a new question of the data sets that back the visualization and get that answer without necessarily having to have a development team uh get involved or an expert get involved to uh give them the answer. Okay. And then uh finally here to um uh
close out I want to talk about agent systems and how we can support uh the development of these uh within a government environment. So um as we've introduced the topic earlier an agent system is simply um an ensemble of AI agents that continuously learn uh your unique data and semantics. Okay. So Genie is a really good example of this. It's a number of AI agents, right? We've talked about how it has access to the data in your data estate via Unity catalog and then you're um uh able to uh ask it questions and then it receives feedback, remembers that and improves itself. Right? So this is something that's a part of the platform. But if an agency wanted to develop their own, we have Mosaic AI which provides a full MLOps capability. Uh so you can uh you have access to the data via Unity catalog. You can uh develop, train and test your own custom models if you want to do that, fine-tune them, uh put them into production with a model registry. So you have that option. You can also use proprietary LLMs. Uh we give you access via a gateway to say an open AI to use their um uh API. You can use open source LLMs that are developed or I'm sorry that are um deployed directly on our control plane. Uh you can fine-tune those. We also support rag architectures. So if you want to use a vector database uh insert your own data into that vector database to quickly serve it to an AI agent. Uh we support that type of architecture as well. So we're giving uh government a significant um MLOps environment in which they can uh hopefully rapidly adopt both AI agents as well as um uh agent systems to improve the delivery of citizen ser uh
services. So closing thoughts. Hopefully uh this was um an interesting introduction to the way in which the data bricks data intelligence platform can be used to um reinvent the way government uses data in AI. Um it wasn't very interactive, right? But now we have the opportunity through a fireside chat uh to listen to um Awesome lead a further discussion on the topic. So Awesome, I'd like to invite you up. Thank you, Eric. Thank you, Eric. [Music] All right, good luck, guys. Just want to start out with asking a quick question. What do you guys think is the most valuable natural resource that we have that we cannot live without? Anyone? Yes, sir. Time. Time. Good try. Okay. Good try. Okay. So you don't need food only time to live. Who said water first? That that's just for you ma'am. So water right? So why water is important? And what happens to the water guys? In other words, when we consume and we use water, what happens to it? And yes, And yes, goes through the water cycle. But what happens to that water? So when you are in the morning waking up brushing your teeth, doing laundry, you know, utensils, industries, manufacturing plants, what are they doing to the water? Yes. Basically making it dirty, right? So it becomes waste water. Awesome. So we all know that on this earth it's 70% water, right? Everybody knows that. Do you know only 3% is fresh water out of 70%. And out of that 3% 2% we cannot consume because it's the ice the icebergs and it's only 1% clean water that we are depending upon. Didn't know that didn't I? So that's why clean water is very important and that's the mission of my organization is clean water for future generations. And one more thought I was just coming here and I'm like huh I met this gentleman we just exchanged the pleasantries and he's like what do you do? I'm like I clean water and he's like what does that have to do with data? Why are you here? Should be in the wastewater conference. Wait a minute. This is exactly what I was wondering. I was like what are you talking about? Isn't this a data and AI summit? And why are we talking about water? Ricky, by the way, this is Ricky Aurora from Minnesota Metro Council. Uh Ricky runs their data and AI practice. And uh today, I know you're disappointed because we won't be you after lunch presenting you with slides, but we'll tell you a story, a story of a data transformation at a public utilities. Um, and we'll walk you through the inception to where we are and see if you know you can actually grab, you know, some some action items and things that you can leverage in your own transformation. Does that sound good? Thank you. Thank you. So Ricky, let's take a step back, man. You need to get give me a chance to get get ready and you So we're going to have fun. Help me first, I guess, tell them a little bit about what does your organization do? I know you started on it. How is it organized? What does it do? And what is your mission? I know you talked about the mission, but um how's it set up? Sure. So again, folks, hello. I'm Ricky Aurora. I'm from Minnesota. So I'm from Metropolitan Council uh Environmental Services. And Metropolitan Council basically it's a regional planning agency and um uh policymaking body and it also provides uh essential services to the citizens of Minnesota. So our customers are citizens like you guys. You know, we we serve about 2.7 million people in about 111 communities. So think about like that waste water that everybody is generating that's coming to our uh water resource recovery facilities to make it clean. Metropolitan Council has five divisions. Uh I'm in the second largest division. The biggest division is the bus and trans uh bus and light rail. of the transit division. Um, mine is environmental services and we operate nine water resource recovery facilities. So, nine treatment plants and they are like all over the Twin City area, 7ount region basically. We also have community development for affordable housing uh and and planning. And the fourth one is transportation services which is contracted services. And fifth one is the support division basically which is all departments like IT centralized you know procurement, HR, legal so and so forth. So they are the they are the support division that supports all the other operating divisions of the council. Um like I said you know we have
uh nine wastewater treatment plants and why is that important? like why uh what does that have to do with data and why am I here and how did data bricks helped us um just to give you a context in our nine wastewater facilities we have uh historians like a SCADA system and I'll explain what that SCADA system is but the data is generated from the wastewater treatment processes like how we treat waste water to make it clean water and how Does does anybody know like how much any guesses like how much waste water do we treat every single day? Any number? Any guesses? No, we meaning in Minnesota. Just focusing on not the nation. How many? Million gallons. A million. Million gallons. Any other guesses a day? Any other guesses? So, we're talking about 27 2.7 million people. We got somebody over there. Yes. 200 million. 200 million in one sold. Yeah. So, where's that star? So, yes. So, we treat 250 million gallons of water a day, Ricky. From uh across um So, there's the star for you. 200 million, right? Very close. So, 250 is the answer. One of our facilities is um 10th largest in the nation. It treats about 180 million gallons a day and then other uh eight utilities combined is about 250 million gallons a day. So the volume is massive. So let's start the journey from because water obviously humanity has been consuming water since humanity has been on the planet and has been working somewhat. What needs to change? What what is your biggest data challenge? what are you trying to change? So, we're trying to transform our uh Metropolitan Council Environmental Services Division to a datadriven decision-making utility so that we can monitor uh the processes that treat this waste water to make it clean water to manage them, optimize them, troubleshoot them and then also at the same time we are trying to make sure that the plant operations who do this operation are also reliable, efficient, cost-effective and compliant. because we uh are also very highly regulated by Minnesota Pollution Control Agency and EPA and whatnot. So there's a lot of uh state and local regulations, air permits, water permits, you name it. So we got to make sure that we don't violate anything because nobody want to be in the newspapers for bad stuff. Good stuff is great, bad stuff is not. And we want to make sure that that doesn't happen. So compliance is a big deal for us as well. Also like I said we have nine different locations where this water comes to. So water moves from these houses through city pipes which are like 6 in to 100 uh 14 ft in diameter. You can run a truck through those pipes and that water comes from like I said different locations to each of these nine plants and that's where this all data is residing right now uh in the historians that collect the data in re in real time meaning we have sensors and it's a real-time data that's being collected through a SCADA system. So that uh for example if something is not right or you know depending upon what operators want to look at they look at this HMI human machine interface the visuals and say oh I need to fix this or I need to change the values. So all of that data that's and our data is huge by the way that's why I'm here. Our data is so across all nine plants we have 120,000 instruments collecting data at 10-second resolution. So every 10 seconds we get 120,000 records. Every 10 seconds there's 1440 minutes in a day, right? times this time 365 times last 10 years of data. Don't even take out your calculators because it will show E. So anyway, so that's the amount of data uh that we have. The other big part is so obviously we want to consolidate all of that data from different historians and we want to integrate with other systems like we have laboratory information management system, we have work order asset management system, we have peopleoft for our financials. all these disparate systems and everybody controls their data. It's my data. You can't have it, right? It's it needs to be just looked at by me. Very common problem. So we when we talk about democratized data this morning, all morning we've been hearing about that. Well, somebody has to unlock that first. So our goal like I said of this transformation is no data left behind. We are going to bring all the data and we started with the SCADA data for these treatment plants. basically and we want to integrate like I said with lab data or other data sets that we have but we also want to empower our end users our stakeholders process engineers plant operations safety specialists compliance you name it all of them so that they have access same access and single version of truth of this data to everybody that's the goal right so empower them with all of this data and then also provide the tools So when we empower them and we have consolidated this data in a single platform lakehouse platform um it just overcomes that data silos the data silos don't exist again we have not done it all so we are on this journey which I'll share where we are at now and how we got here but it's it's a good process that we went through and the last part also is that uh this lakehouse has all the tools that we need for data science advanced analytics. So if once the data is there you have all the tools right now our current state for example is a little bit different which I know is your next question though like what are you doing right now but before we get to that you're a mind reader I am I just captured everything but the idea again is uh to enable the decision-m using these tool sets that are embedded in the platform so we don't have to go anywhere else uh for for data discovery for data exploration all those things that we want near real time access to make these informed decisions. Thank you Ricky. Um so what I heard is the data volume is massive, data variety is massive. It's locked into propriety systems right now which is making accessibility and everything hard and you're trying to decouple and liberate that data. Would that be correct? Yes, exactly. You also mentioned SCA data a couple of times. Do you want to maybe I don't know what SCADA you guys know what SCADA data is? So SCADA is the supervisory control act data acquisition system. So that's the data from the field sensors that coming in real time like every 10 seconds a value is sent from these 120,000 instruments and SCA system is capturing all of that. But then again the question becomes if the data is growing like I said every 10 seconds what do we do with it? How do we analyze it? Right now our biggest challenge is we are only uh with proprietary systems like SQL server and other systems that we have in place we are only able to get to 20% of our data. Our data is we we can't even process all of that data that that's currently coming in. There's no capacity there's no horsepower processing power at all to do any of this stuff. The other uh part of this is besides 20% is that our data is something what we call as it's hydrated meaning if the data value doesn't change like for minute one or 10 second we do not store that value every single time because we will run up run out of storage like every day otherwise. So we only store by exception. But what does that do then? What that does the challenge is well when people ask for data to analyze or you know I want to troubleshoot it's incomplete in their eyes the data is not incomplete right it's just we didn't store every single value every 10 second because we couldn't right so that's the other big bigger challenge for us like how do we take that data the dehydrated data the scala data in this case integrate with like I said other data in the in the plant from laboratory or safety systems. That's one challenge. But then hydrate it back. Meaning fill in those values so that it becomes complete so that a process engineer or operations they don't have to take all of that data first and then hydrate it themselves before they can even start analyzing it. Right? And the third piece is the analyzing of that data right now is is siloed meaning oh I got to take this data now put it in a spreadsheet or run my R script someplace else or Python and then do whatever but then the question becomes well well you have to be here in your office every single morning to press that button to make sure that your scripts are running. I think we all can connect to that. And then once your scripts run then how do you make sure that everybody can access and visualize it on a common centralized place. SharePoint is one but you still have to do a lot of manual effort to get to that point. Right? So that's our current state. Awesome. So we heard why you're changing. We heard the challenge that you have to solve for. We heard your current state. Now let's fast forward a little bit. I know we are already on this journey. What does that future looks like for you on a new data intelligence platform the capabilities that you now have? Sure.
And what does that look like? If you can describe that briefly. So for us it's been awesome because we are in uh initially right now in a state where we are implementing the very flexible lakehouse platform and the medallion architecture. So all these historians and the scala data all that jargon that I just mentioned we are taking all that dehydrated right that exceptionbased data grabbing it from all these historians 100% of the data now not just 20% right 100% of the data putting it into the bronze layer in the metallion architecture in the silver layer that's where we hydrate that data that's where we QC it that's where we make sure that those gaps are there are no longer gaps. It's a continuous time series data every 10 seconds that exists from there. The next is the goal layer. So we go take the data to the goal layer. That's where all the aggregations happen. 10-second data to uh minutely values to hourly, daily, monthly, weekly. It's pre-agregated and it's stored in the goal layer. Again, 100% of the data, not 20%. We solve that problem of it's not exception based anymore. All we need to bring in now is unleash that data and bring all the stakeholders onto the platform so they can start utilizing the data, do the analysis and make sure that they are also uh there's data governance because obviously we are using Unity catalog to make sure that roles and responsibilities and how data is shared that exists and then at the same time we also make sure that the built-in tools so rather than exporting that data from that medallion architecture on the lakehouse platform to a spreadsheet and then doing all of those things why don't you just do all of that Python and R scripting and right on the lakehouse platform and at that time then once it's done you can visualize that on a dashboard it's just all right there in one place one central consolidated place that sounds a lot of fun but if I am a business owner and obviously you know everything requires money and effort what is the business impact what is the return on investment why Would council continue to invest in this? So, this is
all great, right? You have a great vision, but where's the beef? Fun is literally a lot of fun because when you're dealing with this kind of data and then when you take that data and then cater to your end users and it's it becomes an eye openener for them. That's a great feeling. That's a very much fun feeling like satisfied feeling when you see them like whoa how did you get that? I'm like don't worry about that. Just use it. Do what you need to do that you're good at. you do not have to do that other steps but for us the benefit has been the operational efficiencies big time uh increased productivity because again we are not doing that siloed tools anymore so data is just there you start using it innovation so that's when all the communication has started happening back and forth like so I just got an email like I said we are just on this journey so one of the plant managers is like well I want to see this uh chemical optimization and I want to see how that chem chemicals u balance out with the safety because these are dangerous chemicals and there's thresholds so should I lower or you know make those thresholds go go high safety is like oh lower it or go high and she's like no it has to be a balance but guess what we don't even have that data available for them to have this debate because right now it's sitting in those historians and we couldn't bring that because we only we only brought or we could only bring that 20% of the data like I said earlier so these kinds of efficiencies is what now unfortunately we have u nine plants so six plants data massive again it's it's multi- like hundreds of terabytes of data is already in the lakehouse data bricks platform in the medallion architecture but her plant is not yet so I couldn't tell her like hey go to this and we can just you know it's all there so again it's it's coming so operational efficiency is number one the second big thing is the data analytics and the human AI collaboration You know we talked about genie rooms we heard about all of them. So not all process engineers or operations are like people who know R or Python right. So we have we are trying to build those genie rooms in such a way that again simple plain English questions. Hey what's my influent influent flow data from plant XYZ for last three months and gen room just you know shows you a chart and you can export that data. So basically you're saying you're empowering your business users. Empowerment is a big one. But when we put in the AI and then AI identifies those patterns and the have the power to do this multi-terabt of data processing. Um but the domain knowledge exists with the subject matter experts. So when you combine the two that's where the magic is happening. Awesome. Uh I know we got less than one minute. Any parting thoughts?
Um I I think the parting thought is like we started with a very small pilot and when data bricks and asim in this case came to us he's like oh we can help you. I'm like what are you going to help with? He's like well what what's your use case? And I'm like okay here's chemical optimization. He's like no no no give me something big. I'm like okay um here's two use cases. And he's like no you got to dream big. Think big. not just use cases. And I was like getting upset like what the heck does he mean because I'm like okay here's our SQL server you know make it faster. He's like I'm not here to take your SQL server and and make it run faster. No. So then we came up with a journey with him where you know it's not just a tech technological change but it's more like a strategic evolution because like our goal is no data left behind right. So in his case, he helped me see like, oh, you got to dream big, think big, meaning bring all your data. What's holding you back? I'm like, are you sure that you can handle like, you know, so many hundreds of terabytes of data at 10-second resolution? He's like, yeah, I think so. I'm like, I don't think so. He's like, oh yeah, yeah, we can do it. He's like, no. Anyway, so that's the parting thoughts. Like we started with a very small pilot and it was so successful and our leadership is so happy and we haven't even given these tools yet because we are on this journey as we speak. I mean we are processing all this data even right now behind the scenes but there's so much excitement about the the lakehouse architecture the availability of these analytical tools to be able to do data science on them and whatnot that it's just so magical and and people are like looking forward to all of this adoption and they can't wait until we have all the nine plants integrated with all the other uh plant information management systems we have to make that story complete. Thank you Ricky. I know we are out of time so we'll be here if you have any questions but big round of applause for Ricky please. Thank you guys.
Future of data is here.
Millions of gallons, huge data.
Natural language, instant answers.
Incomplete data no more.
AI for data governance.
Unlock data, collaborate freely.
Fragmented data, huge problem.














