The public sector, often perceived as lagging in technological adoption, is undergoing a profound digital transformation. Driven by the urgent need to deliver more with less and meet rising citizen expectations, government agencies are embracing advanced data analytics and artificial intelligence. This shift is not just about technology; it's about fundamentally reimagining how public services are delivered, enhancing security, and fostering unprecedented collaboration.
“These are not use cases of talking about how somebody could click on more ads more effectively, but things that really transform kind of life and society as well.”
- Jude Boyle, Vice President of Public Sector at Databricks
Discover how government agencies are leveraging data and AI to solve critical challenges, from healthcare to defense. Learn about groundbreaking innovations and the strategies driving public sector digital transformation. This session reveals the future of citizen services.
Every day, public sector services touch all of our lives. Roads and bridges connect people, schools educate the next generation. The military and first responders can keep us safe. And taxes provide funding to improve citizen services. But unprecedented challenges and the demand to do more with less is transforming the public sector. Constituents expect interactions with government to be as hyper personalized and responsive as the ones they experience every day with private companies. Governments want to deliver products and services with greater efficiency. And agencies need to upgrade legacy technologies so they can securely share information with one another while shielding sensitive data. So what's holding them back? It starts with the fact that most public sector agencies still rely on legacy technology, where even the simplest changes aren't simple. Outdated data warehouses limit the speed of data, leaving teams waiting longer for critical insights. Meanwhile, government data is growing exponentially with a myriad of data types and sources. And all this data sits captive in silos. The result, responsive data driven, citizen focused government remains frustratingly out of reach. Until now. With the unmatched power of databricks and the Lakehouse for public sector, governments now have one simple, secure and powerful platform for all their data analytics and AI needs. With Lakehouse, you can rapidly build data capabilities that comply with government mandates without getting locked into expensive proprietary technologies. Data teams can comply with even the strictest data policies using fine grained access control and automated data lineage. And now you can easily and securely collaborate in real time with all your commercial and public sector partners. With Databricks Lakehouse for Public Sector for the first time, it's possible to use all your data on your terms with no compromise to deliver on the mission of government. Leading public sector organizations worldwide rely on databricks to deliver a stronger, more secure and mission ready government. Discover the Databricks Lakehouse for Public Sector. Hi.
Hello and welcome to the Databricks Public Sector forum at Data and AI Summit. My name's Jude Boyle. I'm Vice President of Public Sector at Databricks. We're so excited to have you here today. At Databricks, we're fueled by the passion, by the achievements of our customers and we're grateful for your vision, for your commitment, for your investment. The agenda today is focused on sharing best practices and really that's one of the most important things we do at events like this to get the chance to speak with your colleagues, to benefit from their lessons learned. You may have seen in the press. There will be some announcements and we're excited about sharing them tomorrow. Today is all about public sector. I'd say as grateful as we are for the investment for the time you guys took to travel for the time you're investing with us today. That investment wouldn't make sense if you don't come away with some new strategies, with some new partnerships, with some new lessons learned to accelerate bringing the impact of AI to your mission. It's my great pleasure to announce and recognize some of our Data Team finalists. As a leader in artificial intelligence, part of our mission is celebrating the achievement of our public sector customers and amplifying it. There were nearly 300 nominations submitted by organizations across diverse industries and regions in six categories. Each of these organizations have displayed remarkable innovation in their use of data and AI initiatives. We want to help those stories. Here are the public sector finalists in each category. These finalists will be celebrated this evening in our opening conference and the winners in six categories will be announced at that time. So cliffhanger, you got to wait. We have also selected an overall Public Sector Transformation winner that I will announce at the end of forum. Another cliffhanger. I invite these teams to stand as we recognize their achievement. The International Finance Corporation World bank finalists for the Data Transformation Award IFC is a member of the World bank stand to be recognized. Where are you guys as we're speaking about you? Okay, if you're shy, that's all right. The World bank is harnessing the power of big data and AI to address developmental challenges of poverty climate change while making investment decisions guided by responsible and sustainable principles. IFC successfully scaled its AI powered Molina platform using Databricks Lakehouse to accelerate the development of custom machine learning models leveraging faster data processing coupled with GPU computing that scales on demand. IFC designs, trains and runs large language models to analyze massive amounts of data and text. Using natural language processing, molina can analyze 19,000 sentences per minute compared to human readers who read 15 to 20 sentences per minute, a 950x improvement that can reduce the document review time from weeks to days, bringing critical services to those in need. And by leveraging an architecture built on top of Azure and Databricks, Lakehouse IFC unified a diversity of internal and external data sources for both analytics and ML. As a result, the data team completed historical ESG unstructured data with external data such as news company disclosures to develop 10,000 company profiles and expand insights for 180 markets. By leaning on the power of databricks to Lakehouse, IFC can expand its support for emerging markets. By providing this AI powered solution to other investors to also extract actionable insights from unstructured ESG data at scale, contributing to building sustainable emerging markets ifc. Let's just take a second. We'll recognize them. Another nominee, U.S. department of Veterans affairs, is a finalist for the Data Disruptor Award. One may not necessarily think of a social services agency as being known for tech innovation, but that's exactly what the Department of Veteran affairs is doing by embracing the new world of LLMs to reimagine its supply chain operations. The data analytics service within VA's Financial Services center is using generative AI to implement automated scanning and analysis of millions of vendor listings to find the most cost effective equivalent healthcare products using Databricks Lakehouse the journey started with a hackathon to brainstorm new ideas from a broad spectrum of industry collaborators, followed by rapid prototypes leveraging LLMs, a first of its kind in the VA. Large language models help the team to analyze vast amounts of unstructured data, including product descriptions and specifications to identify functionally equivalent items and reduce off contract spending. This transformative project has already achieved significant success. The team has identified automated means to improve spending from 40% to 80%. Through the power of agile data Science, the VA is ensuring superior care to America's veterans by procuring the best available products and services so the budgets can be optimized to spend less on logistics and more on veteran care. VAFSC. The Australian Red Cross Lifeblood is a finalist in Data for Good. The Australian Red Cross Lifeblood knows that saving lives begins with the ease to access biological products so life giving donations can enable life changing outcomes for Australia's most vulnerable. Their performance and analytics team is using data and AI to strengthen engagement and support communities by recruiting new donors, recognizing existing donors and optimizing supply chains. Utilizing Databricks, Lakehouse lifeblood can efficiently run granular and accurate forecasts to understand local dynamics. For example, by processing real time data within their donor centers, they can predict wait times and cancellations, alerting donors of potential appointment delays. They look for opportunities to understand what drives local donations as well as what donors give. Gifts may increase donation frequency through complex segmentation and market attribution models with Databricks Lakehouse Unified Architecture Lifeblood has made it easier to attract more generous donors and strengthen bonds within their community, saving countless lives along the way. Please join me in recognizing the amazing achievements of Lifeblood and all of these thought leaders leveraging data and artificial intelligence to improve critical government services. It is now my distinct pleasure to welcome to the stage one of Databrick's founders. Arsalan, a local boy from dc, cares deeply about public sector to share some thoughts on where we've been and where we're going. Arsalan.
All right. Hey everybody, thanks for making it out. I know some of you traveled from a fair bit of a distance to get here, as Jude mentioned. Arsalan Tavakoli, I'm one of the co founders at Databricks and I run our global field engineering practice. Wanted to spend today doing a couple of things right? One, just basically giving you a sense of from where the company started, what has that lens of our journey been with the public sector practice in general, and then also spending time answering some questions folks may have. I'll try not to kill you with basically PowerPoint, because I'm sure you'll get plenty of that over the course of the next couple of days and today as well. So one, as we look at this, the company has been around for 10 years and a couple of things that I'll mention. When we started right, in 2013, the initial plan was where did we want to focus? It was not the public sector, to be honest. You know, as we went out there, you've got to think about it. We said we wanted to bet on the cloud. That's where it was. And you know, you thought about from a cloud perspective, you know, very much small, medium corporations going down that journey. Now for me, I was super passionate about public sector, as basically Jude said. I grew up in Northern Virginia, kind of in Reston, right down the street from usgs. Kind of obviously exposed to that whole world and tons, you know, knew about all of the data that existed in that world, being able to harness it. And honestly, when you looked at it, you talk about the use cases many of us came from, the founding team came from academic backgrounds. So you thought about how do we actually go not after the simple use cases, but the ones that have a big impact, that can kind of transform society, transform the world, as you might imagine. A lot of that comes out of public sector. I think that the key things though, that intimidated us back then was from a cloud perspective. It's not exactly 10 years ago. The public sector wasn't exactly paving the way to say we want to be the first ones in the cloud. The second was, as you thought about it 10 years ago, going along the data and AI journey, it was still early on and being able to find kind of those capabilities, the desire and the skill set to put that into practice was fairly nascent. And I remember two things happened back then for us. One was of all places that you would imagine, not generally starting would have been the intelligence community, but this happened to be the era that the intelligence community was going towards C2s, right? And looking at what would it mean to be kind of have a cloud for that sector of industry. And there was a couple of change agents across kind of a couple of the public sector organizations. And I remember one of those first meetings was out there was with uscis. Now the weird thing was I live in California now and the federal team was very good. They know when every single one of my family members, my parents birthdays are what is the Persian New Year? What is that? Because they're like, you're probably coming into town, come a day earlier, we have a couple of meetings and clients that we want you to meet. And I remember it was with USCIS that they came in and said, look, we're still early days. 99% of our data sits on prem. But we have all this data and we are just trying to figure out how do we think about people coming in all the different data silos. We have a bunch of data sitting in Oracle that we've been trying to figure out how do we get to the cloud to be able to run analytics on. Is this something that you guys can help us on? We said, fine, let's go down a poc. And we showed them that something that they've been struggling with for basically months and years and under six weeks. We easily moved all of that data into databricks in the cloud. And they were running analytics on it and tackling use cases that they hadn't done before. So that was the start. And from there we saw that there was a bunch of different organizations that were paving the way. Now 10 years, you know, fast forward 10 years later and to be honest, this is incredible to see. Ten years ago in 2013, we had our first data and AI summit. It wasn't even called that, it was called Spark Summit. Back then the number of people that we had could have easily fit into this room. And you just wanted people who could spell the word spark back then. That was the goal. And I look at this now, looking at what are some of the use cases that we've been able to develop. We have whole hundreds of different, basically users across public sector. Everything from, I mean USCIS as mentioned, has continued doing that, working with the CDC on things like disease control, veteran affairs, you know, around basically suicide awareness and prevention. Looking at Los Angeles county and you know, many of the other cities as well, looking at hiring, retention, DoD on supply chain, financial logistics and the like. You know, Jude mentioned a couple of them. It's been amazing to See, and I think that the part that's incredible for me is to say that I think we're just starting to scratch the surface, right? The big, big hurdles, when you think about folks being able to move into, you know, really leveraging generative AI has been a couple fold. It's one, can you move to the cloud? Now I'm not going to sit here and tell you that 99% of public sector's data is in the the cloud, but I think we are past that point where it's just 1% or it's an open question, are we going to put data in the cloud? It's not that taboo. And across basically both federal and kind of state and local government education, you've just seen that main movement go there. I think that the second one is a lot of work has been done around actually setting up the data foundations as well. Can we get our data in there? Can we think about quality? Can we make it consumable in a catalog as well? And so once you have that, it's about where does it go next, right? And so we'll be talking a lot in the next two days about some of those capabilities. But I think the thing that excites me the most is we have now gotten to the point where we talk about, okay, generative AI, what can it do? How do we bring models together, how do we serve, you know, the population and society and look at things as risk and compliance. This group is ready to do that. And still when we look across the spectrum, when you talk about volume of data, kind of mission critical use cases, interesting use cases, ones that could have an impact, I'm hard pressed to find any other vertical that could basically have that breadth and capabilities and potential that this one has. And so with that many of those blockers that historically had been, as I mentioned, the cloud compliance and all those, those and the last one being capabilities, there's such a now breadth of capabilities across both the gsi, the system integrators and the government itself. I think that's what gives me that excitement about what can we achieve in the next 10 years. And so I was telling Jude the other day, getting to come here and talk about this 10 years ago, I didn't know where we would be on this segment. Everybody asked me, did you know that you'd be an X billion dollar company? No, we didn't. We were trying not to die as a company 10 years ago. But now sitting here thinking about what could it be like 10 years from now talking about what this group will have been able to do is something that's truly exciting and honestly humbling because as I said, these are not use cases of talking about how somebody could click on more ads more effectively, but things that really transform kind of life and society as well. So I just wanted to give that quick piece of overview of how I've seen this, you know, journey and then wanted to just take a couple moments to see, you know, any questions that people had. It could be about anything, could be about the product, the company, the journey that we've gone on there, how we think about public sector. You know, who's always wants to be the first person to ask the first question. Go ahead. Thank you. Just question.
Biggest challenge in getting into public sector. Biggest challenge of getting into public sector. Look, I think the challenge from a public sector is oftentimes when you go into an enterprise business, there's usually a very, very clear checklist of what you have to do. It's to say, go into the organization, find the security team, make sure that they give you security approval, go find the economic buyer, figure out what their priorities are and then go do it. That doesn't really exist. You know, when you talk about public sector, the answer is, and having worked with it in the past life, the answer is usually like, well, can you go find this four star general or this three star general or there's this one that might have budget this year or that one. So a lot of it for us, the majority of time spent is how do you get through the Byzantine world of understanding okay, where is the use case, what's useful? I think that that's one. I think that the second thing is we tend to oftentimes focus very much on the technology and the platform, which is important. I happen to think that that's the easiest part. The hardest part is when you start talking about processes and people and change management. Right. And I think if I'm candid, you have pockets of that in organizations. Especially when you look at either the federal government or state and local governments where spread out, there's a lot of people like, look, I've done my job, like I'm keeping doing my job. So how do you find those change agents that say, but look, think about what we could do, how do we basically drive it better? And so it's oftentimes finding that, winning hearts and minds of what is possible and getting to define what the requirements are when we get to that point. The rest of it is driving adoption and success. I think is I don't want to make it to say it's easy, but it's the easier part relative. So figuring out how to navigate that, I think is often something one of the big challenges. Yes, Sir. Testing. All right, Excellent. When it comes to, I've been in public sector for 15 years when it comes to cleanliness of data, like, we know public sector is a little behind the curve, right. So when it comes to AI integration and databricks, how are we looking at moving toward automation of cleanliness and making sure our data is validated when we start using generative AI and things like that? Yeah. So great question. Look, I think I would love to tell you, you made it sound like enterprises have this pristine, great data pool. If you have a population of those enterprises, please send them my way. I would love to go talk to them. Right. I think what I would say on that is everybody has similar challenges, right. And part of it is if you say, look, what's the cleanliness of the data? And then you'll tell me, well, I can't control the upstream sources, where they're coming from. And I'll say, well, can you change them? They're like, actually it's this random COBOL program that generates it. We don't even know where the person who wrote it is. I can't change. So I think that there's a couple of pieces that come about it. I think first of all is defining what are the set of use cases. You'll start like a lot of people start with saying, I want to get clean data and boil the whole ocean. It's really hard. It's going to be a process, it's a journey. Right. Because the short answer is if you start with what is the use cases that you want to drive first, you generally for that have a metric of success and you have some domain experts who can define to you what the role, what are the rules to define what high quality data is. Right. I think if you can do that, the nice part is databricks has a whole bunch of this, right? We focus a lot on generative AI, on, well, are you going to use GPT or MPT or this model? But it starts with the data. So a lot of it we start with how do we get the ingestion to be kind of high quality ones and catalog and curate it so people can find it. That's the big part, right? Because I would also argue you have a whole bunch of data. There's a large portion of your data that exists that there's not actually any use cases for today that you Understand? So what we found is we define, okay, what are the first five or ten that you want to get in? How do we get it cleaned up? People start deriving it. Next we look at where the demand is and go incrementally. If you do that, that we found has been the most manageable piece to go through it. And once you set up that pipeline, you figure out what the first ones are. Usually the first one's the hardest. Cause it's like, what are the rules? Who's consuming it, what are the jobs? What are the SLAs as you get that framework set up? That's one of the big things of databricks, which says, build that pipeline. So as you add subsequent, the level of effort gets basically less and less. Does that make sense? What else? Yes, Sorry, there's like a light directly in my face. So I'm kind of pointing where I see a hand. Thank you. So I feel databricks is really innovating, right? When innovating very quickly and rapidly. So talk about LLMs. You have that. You talk about data governance integrating pretty quickly in the public sector. How do we make sure that people are getting educated quickly enough to be able to adopt them as we speak? What are we doing from that component training and education perspective? It's a great question. I think that there's two aspects of it. There is a enablement aspect and. And then there is a product aspect. Right. What I like to say is that I've got good news and bad news. The good news is that we are arguably one of the most innovative companies you'll find. The bad news is we're one of the most innovative companies you'll find. And so you'll sit on stage tomorrow and basically Thursday, and there'll be tens of announcements of look at all of these capabilities. And on its surface, it seems super overwhelming. You're like, how do we even know to leverage that? So one of the big pieces, and I won't say that we're perfect, but that we've tried to walk through from a product lens, is to say people want to geek out and understand the underlying technology, but how do we make sure that ultimately what they are doing is that they are consuming one platform. I think that that's one of the big challenges, if I'm honest. I think some of the hyperscalers have made is that you have this nice laundry list of things, but then how do you stitch them together? How do you integrate? I don't want you spending time thinking about that. And internally, we have It's a metric we literally report out to the board called cujs which are basically critical user journeys that doesn't pick a feature but it says okay, you want to go out and you want to fire up the platform, ingest some data and basically build a model on it. Show me every step and every click that you do along the way. And we try to figure out okay, how do we simplify that? So that's what you trying to do from the product side to simplify it. I think that there's a second one is to the question that the gentleman asked is usually when we go into a whole organization everybody's like oh I want to know, I want to benchmark you versus somewhere else. I'm like that's great. What's your implementation plan? Right, and so the implementation plan is what is that enablement journey. And we have everything from tons of self service. We put out these MOOCs, these online courses for generative AI and LLMs to all the way to having very, very tailored instructor led courses as well as figuring out how do we set up a center of excellence which says how do we train your power users at that world? How do we make sure that they have key basically chat channels that they can go to and ask questions about it as well. How do we have targeted campaigns to say did you know about this new feature? Here's your workflow that it can make easier. So I don't think that there's a one size fit all but from one, how do we make the product simpler? And honestly, generative AI in many ways is designed to do that. To say your users don't need to understand underneath can they just use natural language to ask their questions? So you'll hear us talk a lot about that tomorrow. But then also a really robust, you know, kind of enablement program that we put in place as well for all of our organizations that we work with. Okay, yes, thank you. So when it comes to working with the public sector, how was your experience dealing with cloud migration, fast speed, gen AI regulation and privacy? The next question is also how does that look like to the end consumer as there is a major, is there any major key success metrics?
Got it. I like that you said last question. You put two questions and your first one had three subparts in it. So now you're just testing my memory. Look, I think the short answer is this is where I would say I know we like to say hey, public sector is very different in this regards. No it's not. Right, so meaning like what you're going to say is, I have an environment that I have kind of a legacy base installed on Prem and I want to move to the cloud. And what you're going to tell me in the move to the cloud is, hey, I want to do it as fast as possible, but you know, remove all risk and do it as cheap as possible. Like, you will be like every other enterprise known to man who will say the same thing. And then you will come to me and you'll say, I'm looking at two different options. Option one, which everybody's telling me is should I just lift and shift? Should I take my data warehouse and move it to a cloud data warehouse and take my data lake and move it to a data lake and then figure out how to modernize later or lift, modernize and shift. And so the answer that I would give you is never do the first. Right? Because what will happen is you're gonna say, look, it's the same. I'm gonna move from one data warehouse to another. No it's not. It's gonna take you longer than you thought. It's gonna be more expensive and harder and once you get there, the rest of the organization will not have the budget or the appetite for you to do the same second modernization step and you'll be stuck without actually changing any of basically the business capabilities that you needed. So mostly organizations come to us and say I want to lift, modernize and shift and land on which the answer is landing on one enterprise data platform. What we spend a lot of time doing is that we both ourselves working alongside of a lot of the SI partners, have a pretty detailed playbook, pretty much any system that you can think of, it's saying either I was on Exadata or I was on Netezza or I was on Teradata or I was on Cloudera or I have my postgres databases, the like of it, we've done a bunch of those. How do we help you pick that up and migrate it into the cloud along our SI partners? And having done that now thousands of time, we have examples here is kind of what are some of the pitfalls of doing it. The nice part is generative AI has a bunch of really amazing technology as well. Because many of the code that you're getting, you're going to have some old school code. Generative AI can look at it and tell you this is what the code did. It can also easily basically transpile and say here's how you move it from your, you know, random pl, SQL language over Here to move it to a different ANSI SQL dialect so it does those pieces for you. That's one of the big things that folks ask us is to say both how do we do that migration? And we've seen enough of it. How do you speed up my time to get getting there? The big challenge I have found is when people start like when they do that migration, they don't, I know it sound like a broken record. They don't think about what is my process perspective or what is the people so great I moved everybody there, what is my path to getting the users enabled on the new environment, how do I make sure that the data I'm putting there is high quality and consumable. And so usually on migrations rather than the technology to technology, that's where we're spending a bunch of time time on and we've got a bunch of examples that people have taken and say in less than six months moved massive on premise systems into the cloud. And the advantage of being one that's later to the game of doing it is that you have a whole body of folks who've gone before you and kind of hit all of the paper cut and hurdles that you think and ideally we can help you make sure that you don't hit those. How did I do the three subparts in the second question? We're okay. Okay. Do we have time for one more? Okay, sorry. Well, I'll be around for the rest of the day but again, a big thank you for this group. I think some of the use cases you're working on are amazing and look forward to continuing to partner with you on it. Thank you. Thanks so much. Arsalan. We're so excited to have him involved with our customers. We've got a great panel coming up. I got to use the clicker, huh? Okay. Can I welcome to the stage Aaron Kenworthy, Young Bang and Nick Lanham. Why don't you guys join me in welcoming them. This is going to be a real treat. I'll let you guys introduce yourselves.
Okay. Excellent. All right, it's on. Thank you, Jude. So you know Arsenal, I made the comment about the public sector and you know, kind of the impression of not being forward looking. But on stage with me today are two executives from the Department of Defense that have done everything but that they are really, truly leading the way with their innovation and what they're doing to change the way the Department of Defense goes to their missions and handles data. So what I'll do is just quickly, you know, go down the panel. Mr. Bang, why don't you tell us a little bit about yourself and what you're doing with the data mesh initiatives that you have for Enterprise at the Army. Sure. Thanks, Aaron. So my name is Yong Bang. I'm the. You can see the title, it's really long and people are like, what does that really mean? So really we develop or buy or integrate all the weapon systems for the army or all the computer systems, or all the data systems or the cyber systems. Right. So offensive defense of all that. And so some of the biggest things that we are really driving at is that we're trying to abstract hardware and software and abstract data from software as well. And I'm sure we could talk a little bit more about those type of things, but we do that for multiple reasons. But around the data side, data is tightly coupled with software right now. But we gotta get the data problem right. Anyone who's done any type of system implementation knows that data migration is always the biggest issue. Right? And so if, and someone alluded to the data maintenance problem and the data wrangling that we all have to do, well, until we get that fixed, we're not going to get the AI at scale that we need. And so for us, one of the things that we're doing right now for us is really table stakes is to try to square away our data enterprise and flatten it and set, simplify it. And so for us, we're really driving this whole motion of a data mesh. We're not adopters, we're driving it across the board and we're actually tweaking a lot of the concepts on it. So we think again that will help us be accelerated to enable better, faster solutions at scale to include artificial intelligence. I appreciate that and thank you for explaining your title because that was a. I was not going to memorize that and then try to do it right. So you saved me. Mr. Vanham, a little bit about you. Yeah. Good afternoon, everybody. I'm sitting in for Greg Ludle today. I'm one of his deputies. I work in the Chief Digital and Artificial Intelligence office and I lead up a group called the Enterprise Platforms and Capabilities branch. I'm an Advanta co founder, really excited to be here today with you again. We're going to talk to you a lot about how databricks is part of our platform, but I want you to know it's part of our central nervous system that runs the entire environment for us. So connecting and supporting the data mesh that Mr. Bangan mentioned, how we're doing inbound and Outbound data connections and how we're managing them, connecting them to our catalog like Arsalan mentioned. We're going to talk to you about that today, but we're really, really excited to be here and appreciate the opportunity to support and likewise, and just for the audience today, maybe you could just give us a little bit of explanation what Advanta is, right, and what it's being done or used for other dods consuming it. We'd love to know the background on it. Okay, great, I'd love to tell you about it. So I think some of the questions here is a perfect example. Is anybody in the audience know just how to raise your hands? Who knows what Advanta is? Have you heard about Advanta? Well, that's pretty cool. So I think if you think about this like, you know, it's very humbling because we started about five years ago. So as Arsalan was going through his opening comments, it made me think back of when we were running around the Pentagon asking everybody like, hey, have you heard about Avanta? Do you want to use Avanta? And really what our platform has become? So what is it, right? What is Avanta? We're basically an amalgamation of COTS products integrated in a way to help support data automation, data collection, data analytics at scale and operate on every classification environment in the Department of Defense. So if you think about some things that are really easy in the, you know, like kind of the commercial world, right? I can go to a, get any type of notebook that's out there and I can write Python, SQL, Scala, I can get access to data. Well, in DoD we have thousands of analysts that can't do that, right? So that import statement that you can write in databricks where I can just say import tensorflow and then import scikit and then import anything else, that's a really big deal for us in DoD because just access to those enterprise tools is a challenge. That's one of the things that Evana is trying to help break down is access to enterprise tools, really powerful tools connected to enterprise data, which we've spent the better part of, I don't know guys, about five, six plus years here of just working to connect to all the different enterprise data sources across DoD. So if you guys are familiar with the Deputies creating Data Advantage Memo, that's a really important memo, but it really talks about senior leadership using authoritative data sources to make data driven decisions during their senior governance forums. And that's honestly, it's really awesome that we've had the opportunity to support that Use case. It's one of our marquee use cases along with supporting the DoD audit. And really what those are all about is what kind of Dr. Martell tells us to do, which is really get the data layer right. It's for the first time ever, connecting to every enterprise system that we can in that particular use case or in that area. So think for the audit, one of the challenges we had in DoD was we have 20 to 30 plus general ledger financial accounting systems out of thousands of systems that operate in the DoD. So you can imagine every one of them has a team that's focused on, on their engineering and that's who you've got to work with and talk to and actually create a connection. We're systematically trying to create a connection with pretty much every one of those systems that is in that use case bundle for us to be able to support. And then on top of that, our users come into Advanta and as long as they can access our tools, right, they can actually log in and use tools like databricks to start exploring that data. Right. And integrating in ways that you just never have been able to explore or integrate in the past. And I can tell you from being a user in databricks, I remember my story is that I like to tell people, I worked in the basement of the Pentagon for some time for a group called Cape and we were doing some amazing use cases there. We were coding, but something as simple as getting access to Python, that's like a software form per machine that we had to submit. It's just somewhat time consuming to get those approvals. I remember when I joined Greg's team and he showed me for the first time ever for the audit, it's like, hey, we have all 30 plus general ledger accounting systems coming into a single place for time. The first, first time ever to support this thing called the Audit. Do you want to join? I was like, yeah, this sounds incredible, let's do this. But I remember when I joined, what he didn't tell me is we didn't have all the tools that we wanted to come over to start coding and developing at that time. And that's okay because we were really working on building out the platform, kind of like Arsalan mentioned earlier. At that point, we really tried to evolve as a platform and add all the tools and products and integrations to help support all of our users. But I remember the day when James Doswell and our team came forward and they talked to us about, hey, we got this new tool called Databricks that We integrated into the platform. You should definitely check it out. And I remember just thinking, I was like, wow, this is one of the tools that we really need to help us explore and move fast to actually have access to all these libraries and packages that we haven't had in the past. We just use it so much. I think we're up to about 3,000 or 5,000 users now. So I know that might not sound like a lot to everybody who's in the millions of users world, but we're really proud that we went from about 3,000 to 5,000 users in total as a platform to about 72,000 now. And we support senior leadership across the department. We do everything from data integration, aiml, data cataloging. And then I'll talk to you a little bit more about our data integration operations layer, which is the core that actually automatically moves files back and forth. But hopefully that answers the question or helps. It does. And you've had infection, phenomenal growth. Right. It's just been incredible to watch. And what is interesting and when we were preparing last week and talking about it, is that these aren't competing platforms. So, you know, Mr. Bang, I was really impressed with the way that you kind of perceived, hey, here's what we're doing in the Army. Right. Which is impressive in itself, but also that it complements, you know, an initiative out of the CDA like Advanta. I'd love to have your thoughts on how that works and what's your perception? Yeah, I mean, so. So for me, I don't think things are a binary thing. It's not this or that. A lot of times it's a hybrid environment. And when you think about Abana, we use it as well in the Army. And just like Nick had said, it aggregates a lot of different data sets, especially on the financial and audit side of things, which is awesome. And when you're connected and have robust pipes, robust storage and robust processing, you have the advantage of that. But army, we tend to have a density and tend to be disconnected a lot of times. And so for me, when we think about that, it's nice when you're at the Enterprise or at the installation, you can access everything that Nick and Greg are doing with the ABANDA platform. But a lot of times our service members aren't exposed to that or don't have access to it. And so we have to think a little more tactically. Right. And it's even different from the Navy. Right. Because Navy has a carrier strike group with multiple ships or boats. I don't know what you call them. And they have an infrastructure. They have storage, they have processing, they have comms as well. And so they can actually access that. But on the army side, it's a totally different scenario. And we have our soldiers are our platforms. And so we have millions of folks out there that we have to that won't actually have access to all that. So again, where we're thinking through and looking at how do we leverage capabilities at the enterprise level, but how do we also augment that at the tactical side? And that's why we think, again, it's not a binary solution. It's not just a platform. When you're exposed to the enterprise, how do we get that in a mesh environment for our tactical community? And so think about the complexities that. Right, that mesh and the data mesh, when you think about the data product, when it's localized here, that localized version is actually enterprise for that unit or that brigade or that division. But when it connects back up, it's got to be actually local versus the enterprise. And we want to consume a lot of the services and that type of information. And so we're really thinking through what scenarios do you aggregate or centralize? Where do we actually decentralize and do federated query more often than than anything else? And then how do we look at data and do we persist data? Like everyone likes to aggregate it and persist it in plus one architecture, we can't afford to do that. And so for us, it's really about how do you balance the two, how do you work it together, how do you look at the mission threads and really think through the differences? And arguably, in a wartime environment or scenario, when you call for fire, I don't need all the information. I just need a little bit of information to say I need to do XYZ and get me the most relevant information. I can't worry about stale data or this duplicate data or this information is not accurate. I just need to know what I need to know so I can make a decision, like, very timely. And so the scenario that you talked about, again, a lot of people think about the enterprise, but the Army's the most dense population with the most information across the DoD. And if you want to help us with hard problems, come help us with the army to figure out how do we work this. And it's not an either or, which I thought was really impressive. Right. And very often the DoD can get accused of not, you know, sharing toys in the sandbox and collaborating. And yet this is a situation where, you know, many of the initiatives you're doing to make the pyramid flat. You know, Advanta can be an and or a supplement or a great way to start and take. So, Mr. Williams, is there something you'd like to share with others or how do people get started? The groups that aren't using Advanta yet? What is it for them? Yeah, well, so I think at the kind of core, it's access to enterprise data, because we are Policy Memo, we host the federated catalog for the dod, which I think is a really big use case just in itself because it represents one of the first times ever that we actually have a goal to connect to many different enterprise systems across the DoD. And really, like, you know, when that J4 question comes up about logistics across DoD, it's not just one system, right? We're connecting to many different logistics systems across DoD and putting it into the catalog. And we're building that automated pipeline once there, per system that the army, the Navy, the Air Force, everybody can leverage versus them having to go back and recreate those same pipelines over and over again, and us having to Procure the same Atos for those same products and tools like databricks, right, to use. So what's in it for them to start is access to that enterprise data catalog. And then it's really up to you, really, if you want to use it more for data management and analytics at the App layer and create automations that feed analytics, that feed senior leader dashboards, and use those there. We support lots of those different types of use cases. Cody and Aaron, who are leaders on our data and AI management team that actually create the data integration layer, that's one of our most popular use cases that we support. It's actually what automatically moves data around within the Department of Defense. And if you get an idea of how massive that is, we're talking about, you know, every time we create a new data connection, we have a process for setting up what we call a memorandum of agreement. And, you know, typically like an sla, right, with that data source that tells us, hey, this data is of this data type. It should not include this kind of information in it, or it should include this kind of information in it. Please put these IAM policies around it and put these security roles around it. We turn all of that into automation and rules that are managed in a, you know, like, automated way using our data integration layer. And then I think one of the most critical things as part of that is we have a structured way of moving data around the platform and both Inbound and outbound. So if you've heard of a medallion architecture, bronze, silver, gold, right. That means a lot to us. We, you know, pretty much are, you know, creating those new data connections that land in our bronze zone. Right. And as we do more data automations to review and apply data quality to all the data pipelines that come in and a common data model to that if we can. Right. We're moving it into the silver zone and then from there we're trying to integrate multiple data sources that have not been integrated before. And to do that you need powerful tools like databricks with a lot of libraries and packages to let analysts get in there and explore and discover what actually matters for senior leadership. Right. Let the data start to tell us more about what is it that we should be looking at for from an analytical perspective versus kind of just taking the hey, I think this is an issue and let's go try to find out if we can or not disprove or disprove that. What we're doing here is really being able to see ourselves in a way, kind of using a single source of truth connected to our authoritative data sources and leveraging the ability for us to let the data tell us kind of where some of the issues are. So I think final bit of that would be that automation layer allows you as a user or a command to really move data in both inbound to Advanta and use our enterprise tools that we have. But you can also federate it outbound. And to Mr. Bang's point, that's really what's part about the data mesh in DoD. And what's so critical about it is for us to be able to federate data outbound. And that's a really big deal for us as we go more and more into that enterprise. I mean we do a lot right now on cloud based connections for kind of S3 to S3 connections. But we're actually in the process of integrating an API layer into right now in Nipper. So that's going to be a new world for us. Directly requested by Dr. Martell for it's one of our main things that we added into our environment. But that API layer is coming online in the next couple months. It's going to be our first kind of test case where we're actually letting folks come in, in a dev hub kind of environment, kind of create their own API with their own credentials and put that into your application and your dev process. So we're pretty excited about that. But for what's in it for You I'd say that the core, it's the access to enterprise data, the federated catalog, but then it's also access to our enterprise tools in an atoed environment that you could use to kind of shape how you want to support your use case over. I appreciate it. And let's go back in time a few years, right? Databricks is new on the horizon. It's this exciting opportunity. But the fear and the lessons we've learned about vendor lock, right? How did that shape your perception? And we'll start with Mr. Bae just on your approach going forward with Data Mesh. How did it influence you? What are you doing about it? How did databricks kind of fall into that and affect it? Yeah, so before I came back into the army, I used to be on the commercial side working in consulting companies or startups or product development companies. And a lot of companies take an approach where they want to vertically integrate for efficiencies and all that type of stuff. But let's face it, it's for revenue. And when you think about that in the context of DoD and the army, we have nothing but vertically integrated stovepipe systems. And go back to what we just talked about earlier when I was starting out. We're trying to abstract hardware, software and data because they are so vertically integrated and we are vendor locked and we can't even get our data out right. So this abstraction layer getting to a single source, like Nick is doing with the Avanta component, helps us create abstraction layers. And for me, you know, I almost made a career about figuring out interchangeability and interoperability around data and systems, right on the consulting side of things. And you know, I'm going to go back a little bit. About seven years ago, me and my team were working on something creating actually back then what's now known as a data platform. But we actually used Apache Spark and we're like, oh, this is amazing and blah blah, blah, this. And we kind of looked into who was the biggest contributor code to Spark. And it was this company called Databricks was really small, 2, 3 years old at that time. They were the most prolific contributors to the code set. And then we kind of did a little more research and we found they had optimized, obviously not part of the Spark portions, but they had optimized performance and everything on that. And we're like, it'd be great if we can just plug and play that back in. And you know, having worked in this industry for 30 years, nothing's ever that easy ever, right. So when we actually Dropped in and I'm dating myself dropped in dbr and instead of Spark, it actually worked without manual intervention, without us adjusting everything onto our platform, we're like, oh my God, this is amazing. And so that was really when we really kind of jumped on the bandwagon, really embraced databricks and their approach on open architecture. Right. And they do it so you can actually, even the competitors can actually consume in different components of everything that they have. And for me, that really helped me to shape my thinking. So now shift it back into the army and DoD. We were talking earlier about MOSA. MOSA is the modular open systems architecture the DoD is trying to push. That's been around for years. It hasn't worked because we took a very standard based approach which really became a checklist, a compliance checklist. But it never really enforced interoperability or interchangeability. And so what we're trying to do differently in the context of Data Mesh is we're actually building out a reference architecture, a reference implementation that we're going to refine the reference architecture and then we're going to design, develop design patterns specifically to ensure. Right. MOSA plug and play. And we're really driving something called torque traceability, observability, replaceability. And how do you get there is automated consumption. Right. So you can be replaceable. But if I have to spend, you know, 30 FTEs for five months just to convert or migrate over a data set, it's not really consumable for me. And so when you think about all those things, it's really because of we want to drive open, open architecture like y'. All, we want to really think about how do we automate that so we can actually enforce it in a design pattern. Right, right. So we give guidance and some flexibility, but it's not so broad interpretation that we're not interoperable or interchangeable. No, that's great, Mr. Whan. Same thing. Anything with the vendor lock kind of approach that shaped Advanta in your solution? Yeah, I mean there's so much to that question. So I mean one of our core tenants is no vendor lock in. Right. So I think that's something we're really, you know, Greg's very, he's driven that into our DNA as a lead of the platform and for us it's really important. So as we add new vendor partners and new tech to our stack, we're really trying to make sure that we can integrate those products with other products and services in our environment. So think about integrations with our catalog, with our data Operations layer that I mentioned earlier. With our application layers we have a lot of vendor Partners, Tableau Qlik, C3AI Virtualytics. We have a lot of different builders leveraging not only databricks but the integrations that exist with all of those other tools for us. But when we make selections about what tools go into Advanta, we're very, very, we try to be very strategic anyway about what tools and products we're adding into the platform. Because those products are, you know, they need to be integratable and we need to make sure that we have the ability to move and shift if we have to, if we're concerned about vendor lock in. So for us it's a very, very important concept for us. And I mean so much so to where we have made pretty massive and drastic changes to our platform as we've had to make shifts as we've grown. Right. And to support the scale that we're getting at. So I want everybody to kind of know that. But I think on the same sheet of music, one of the things that we do here in the VANA platform as part of CDAO is every three months we're evaluating as part of our technical roadmap status meeting what products and new libraries and packages and capabilities are being requested across the Department of Defense or the folks that we support. Right. And what we're doing is making strategic decisions there to say, okay, is that a new product that actually aligns with our strategic roadmap? Right. Is it something that Dr. Martel or Ms. Palmieri or Greg wants us to add to fill in a capability gap? Is there a new feature that maybe Databricks is going to add, maybe a software as a service offering, maybe in the future those kind of things. So what we're do trying to to do there is figure out are those new features capability or requests, are they integratable with all the other products and services we have? And then we make strategic decisions about do we add them into the platform or not. And there's a whole process with that. But no vendor lock in is at the heart and soul of every one of those conversations and certainly in the decision making process. And I think it's the real differentiator right from the approach that you're doing. If you look at, let's not repeat things and try to get different results and you're taking that approach, you're putting the architecture in place. It's a significant change in just like culture like we were talking about this morning and the way that you're doing it, and I think it's a sign of your success, right? And it's showing that it's working. Just like we go back in time now, we can look forward, right? So Arsalan was mentioning, you know, we're going to be making all of announcements, and certainly generative AI marginal model formats are, you know, all the rage. Surely the DoD has not adopted that, right? Everyone thinks you're five years behind. So is this something that's on your plate, you know, Mr. Bang? Are you guys even looking at something like that, or is it too far out?
Yeah, I mean, so there has been a lot of talk about large language models, and I always say, great, we should probably look at small language models, because performance of those. Right, has been incredible for the past year and a half, right? The performance of those is just as good as the LLMs, and arguably it's more applicable for us on the DD, right, with the different terminology and training data sets that we have, that could be a little bit more focused for the domain. The other part about generative AI, I mean, I love this one, right? So if y' all humor me, the Army's been doing artificial intelligence just like everyone else in pockets. We're trying to do things so we can scale this and make it bigger, better, faster, stronger. But, you know, let's take an example, like Covid. Everyone's been touched by Covid, so I want to talk about that. And the army does some incredible work, but we are the worst at actually publicizing some of the work that we've done. And so if you think about basic things like, oh, I don't know, vaccinations and test kits wasn't actually hhs. No offense to the VA or anyone else that did that. It was all army, right? Operation Warp Speed, but specifically the acquisition, the vaccines, all that type of stuff. And the Army's still actually doing it. We're trying to transition it over to hhs. But if you think about that, that's. I just want to get a shout out to the Army. But when you think about that, remember a little while ago, a year and a half ago, Omicron came out. The variants that came out before that, and the shots and the virus and the. I'm sorry, the shots that we got for that weren't affected on Omicron. And so this is where I actually want to talk a little bit more about what the army has done. We've actually used generative AI to do that. So if you think about why Covid and Omicron was Such a hit on the industry is because it had mutated so fast that the antibodies or the vaccines that we had were not effective against that. And so we actually looked at that problem. And yes, it was the army that actually did this in conjunction with Lawrence Livermore lab right up the road here. And so we did a combination of simulations around that simulations as well as generative AI to look at co optimizing how those type of things will work. So if you think about that, because the virus had mutated so quickly, we didn't actually have much information around the signature and the mutation. And so how does things work? Just like in cyber, we have a signature, we have a right, we do, we have a remediation against that vaccine. Same way, if you have variations and you don't have a lot of information, how are you going to really accelerate some type of development of a vaccine against that? And that vaccine development usually takes about 15 years, right through the FDA process, clinical trials and all those type of things. And so what we said was we have to take a different approach. We have to look at simulations as it relates to the characteristics of the virus, characteristics of antibodies that we have against some of the ones that were actually effective, look at how they actually bond. Right. And so while we actually, what we call co optimization between the simulation and generative AI to look at all the possible variants and then optimize the antibodies, right. To bond against this new mutation of virus. Right. And then also have it backwards compatible. So this new virus, I'm sorry, the antibody actually worked against Omicron and the previous ones, where the other ones only worked against the previous ones, but not against Omnicron. And so think about the possibilities or how hard that would be. We couldn't have done it without generative AI. Right. And so that, that really accelerated the timeline. And when this was discovered, I don't know, in the December ish timeframe, we had something out in January for a mutation without a known antibody, without any known data around that mutation. And that's actually what everyone's gotten a shot on, right, to really address that. So again, I would say the army has really embraced AI, generative AI, even different types of language models, not just large language models. But again, we want to really put this in context and really do this at scale. And so that's why the data component is really table stakes for us. That's great. Mr. Manham. How about you guys? Are you guys thinking about using it? Oh man, there's so much to say here. So we only have about Two minutes left. So just caveat that with everything else I was saying. But. So one thing I think we could say is there's insane demand for LLM. Specifically, like Mr. Bank had mentioned, we're seeing that across the board. Our leadership is obviously concerned with folks getting kind of enamored with the answer and not kind of being able to automatically and have a scalable way to check the results of that. So that's what we're thinking about constantly. We're very excited about databricks. Dolly. I think some of the things that I don't know if the room knows here, but we've installed it and basically deployed it in nipper in our nipper environment. And we're working right now to actually see if we can get it installed in sipper, which gives us the abilities to start to actually not only leverage the generalized LLM but also have an actual opportunity to start to train that on corpuses of DoD data in that classification zone, which is a really big kind of, I would say a big cutting edge capability that folks might not be aware of. But yes, we are very much interested in leveraging it. We appreciate databricks and, and innovating that and partnering with our team to help deploy it. And again, the demand there is just so much right now. So it's just trying to keep up with that. No, we appreciate it and that's exciting to hear. So the minute that we have left kind of an open ended question for either one of you. There's an entire range of experience here, some people that are new with databricks as well. You guys have done incredible fast paced type of a dialogue and implementations of this. What advice would you give people, whether it's a lesson learned or where to start that would really help them on their journey? Sir, you start. Yeah, I mean I think at the end of the day, being a solution architect myself. Right. We jumped to the technical portion. I think there were some discussions earlier this morning and then someone alluded to about people, processes and those type of things. I think governance is a component, but I think a lot of times as technologists we jump to this technical solution or the shiny bauble. I think a lot of times we have to kind of step back up just like we're abstracting things up. Step back up and truly identify what are the questions we're trying to solve in the military. It's really what are the decisions we have to make that should drive what data that you need? We don't need everything. Right in my mind. And so I Think we have. We talked a little bit about this. We have gotten so good with technologies that we really created this inverted pyramid where we have so much. The more senior you get, you have so much information, so much data. It's data overload. And that's again, because we focus on the technology aspect. I love to really re flip that over. So it really becomes a pyramid again. So the decision makers at the top, right, can really get the information that they need. They have access to everything else, but they don't actually need all that. And so really, how do we look at that? How do we look at what are the critical answers that you need or the decisions or the questions that you need answer and use that to really streamline or simplify our architecture. Well, I can't thank you both enough. What you've done for the military, for the Department of Defense with databricks. We're honored to have you guys as business partners and thank you for your time today. Appreciate it. What an exciting point of view from our thought leaders in the Department of Defense. We have an exciting conversation now with a sled point of view. Are we doing. We're doing this now. Can I welcome to the stage our friends from California Department of Public Health, John Russell, Michael Powell and Ajahn. Thank you very much. Join me in welcoming them.
Thank you. Super bright. I was going to try to come up with something slick to say about the Department of the army and the difference in the technology that was used when I was in the army to what they're doing now, but that was only 34 years ago and they might still be using that technology. So I don't want to get anybody in trouble. So thank you for joining us today. My name is Michael Powell. I'm the chief with the registry and assessment section with the Immunization branch. For the full title, that's with the Division of Communicable Disease Control, center for Infectious Disease, California Department of Public Health. I'm joined today by John Roselle, the CIO for California Department of Public Health and Ajali Sen with Accenture, who plays World of Warcraft. Anybody in the room? Don't be shy. There's an achievement in World of Warcraft called what a long, strange Trip. What a long, strange trip it's been from the time you start to the time you finish. It's no less than one year to get that achievement. I'm going to take you on a journey today of our long, strange trip and what our trip is going to be. My long, strange trip began In January of 2020, when I was at one of my favorite conferences in the world. It was a budget meeting at the cdc. And during the conference on day one, we went to many different rooms, did our budget meetings. We were told how much money we had, what we were allowed to spend it on, or how much money they were taking away in some circumstances. Day two of the conference, we show up and the conference rooms are closed. Nope, you can't use this room because our epidemiologists are using it. But that's all we heard. So then they kept moving us from room to room to room. By day three of the conference, novel coronavirus was all over all of the rooms. Later that month, WHO declared, for only the sixth time in history, a global health emergency, which you all are very fond of now as Covid. So I'm going to take you on our Covid journey, which led to much more of a journey shortly after the who. Where's my clicker? Oh, there we go. It's your responsibility now. I really don't use the slides anyway, but. But I'm going to put this slide up because this kind of shows our journey. I'm not going to read the slide as it is, but I'm going to kind of hit some of the high points for you. So after the WHO declared their global health emergency, shortly after, obviously, for those of that are in California, know that In March of 2020, we had our shutdown. Everything was just stopped. You couldn't go to a restaurant, you couldn't go to a bar, you couldn't go anywhere. It was even hard to go to Sam's. And if you went to Sam's club, you bought $700 worth of groceries. If you didn't, maybe I'm alone, because I did. But we knew early on right when that happened, what was going to happen. Not because we all knew what Covid meant or what Covid was going to be, but some of us had experienced H1N1 in 2009. And in California, we had a specific potential Testis outbreak in 2010, a measles outbreak at Disneyland in 2014. So we kind of knew what the pattern was. And I was a scout leader as well. And I used to tell my scouts and the parents, I will give you 100% to scouts until there's an outbreak. Once there's an outbreak, I have to go away. I don't know where I'm going to go or what I'm going to do, but I will go away. Man, I had no idea what it meant when there was A pandemic because we were working 18 hours a day, seven days a week, and it was insane. But we didn't have a vaccination yet. And I work in vaccinations. Before we had a vaccination, we had to have contact tracing, we had to know who had the disease, we had to do testing, we had to do disease reporting. And we had systems that for a long time were limited in funding. We didn't know, you know, we really didn't know what our capabilities were because we didn't have a lot of data coming in. So for 2020, that was our life, was learning how to get the disease results in. How are we going to analyze the data? How are we going to try to predict what's going to happen? The State pulled in 10,000 redirected staff to help with this effort, with testing, with getting our systems ready for immunizations, even though we didn't have them yet. And we also knew that our systems were old in some cases. And we knew that we had to look at ways to improve what we were going to do. In December of 2020, we got our first vaccination. December 14, 2020, to be specific. And the reason I remember that date is because we were all watching, because we knew once we got that one vaccination, just hundreds of thousands of vaccinations were going to come in at once. That wasn't true either. We got six. But the good news is they all processed successfully, so we thought we were on a good path. In January of 2021, we realized very quickly that we did not have the data systems in place that we needed to know what was really happening. We knew that when we started getting vaccination data, we had to look at vaccine equity. How are we serving our populations? Who's getting vaccinated? Where are they getting vaccinated? Where are they not getting vaccinated? These are all the things we really needed to know. And this is another part of the journey. At like 4 or 2:30 in the afternoon on Friday, I got a call, said, hey Mike, we have to set up a data system for Covid. I worked tirelessly for 14 hours with a whole group from the Department of Technology. And by five o' clock the next morning, we stood up a data system in a few hours ready for everyone to use. And was that my walk off music? Yeah, am I done? Am I taking too long? And so now I lost my train of thought. But anyway, we knew that we had to get all these things in place, so we knew we had to be data driven. And we were not a data driven system at the time. Over the next couple of months, we did start to get more data and more data and more data. And in the spring of 2021, it got so bad that we essentially tore our entire system down and rebuilt it several times to try to get it to handle the data. At one point in California, we were getting 4 million messages a day for vaccinations coming in, people querying the system to see who got vaccinated. Then just in general, we have a process that goes through and finds possible duplicates in the system so we can do patient matching. All these things just overloaded the system. We ended up moving our system. Not to be too specific about servers, because I know this isn't a hardware course or anything like that. We went from 18 servers to 96 servers to handle the volume of data data over just a few weeks to get that system up and running. In June of 2021, the next dramatic step in our journey happened. Two things happened in June of 2021. One was the state reopened and two, California developed a digital vaccine record. Or at the time, it was the digital COVID vaccine record. This was our version of consumer access that would allow anybody in California who was vaccinated in California to go to our website, get their record, have it on their phone with a 3D or QR code, and they could take it to their favorite bar, not that I did that or anything, to their favorite bar, restaurant, use it as proof that they've been vaccinated. This was a huge success. And it was so successful that we ended up working with Washington, Oregon, Virginia. I'm looking at Amanda in the back because she was the project manager. Michigan, several states we worked with to adopt our code so they could have their own version of Consumer Access. It was all built on the smart health card framework and was globally usable. Many international airports also used the same Covid or the same codes. Then in January of this year, the digital COVID vaccine record became the digital vaccine record. I'm just throwing that out there because we're really, really proud that that happen. So now you can get all your vaccinations on your phone as a consumer if you're vaccinated in California. Then after June of 2021, we spoke a little bit about Omicron and there was also a couple of other variations of the vaccine. Since then, we've also deployed the bivalent vaccine. We've submitted 680 million vaccinations across 38 vaccine groups have been submitted to the registry. 70 million Covid doses have been administered in California, which is really big. That's a great accomplishment. We had a goal from the CDC to vaccinate 70% of the population by July 4th of 2021. We hit that in June of 2021, which is why we were able to reopen California. But we couldn't have done this without a lot of help, a lot of assistance from others. ITSD was huge in everything that we did. Accenture was huge in everything that we did. And with that I'm going to turn actually I completely missed my number slide that I was just talking about. I told you, I really don't watch the slides, but these slides are available for you and I'm going to turn it over to Ajali now and she's going to walk you through. I'm actually going to give you the clicker. Sure. Thanks, Micah. Okay, so one of the things that Michael mentioned is really talking about the journey of a transformation of CDPH into a data driven organization. And I heard this mentioned before. But where our success truly stands out is how data has been integrated into the very core processes that CDPH undertakes, not just at a strategic level, but in a day to day operational level. And what this enables is a data driven public health policymaking or data driven actions that enhance public health. And this can be really brought to life through a day in the life of somebody who works in the California Department of Public Health Health. The person who I chose, Jennifer, she's fictitious. But everything that we have represented here is true. And really the purpose of this representation of day in the life is to show that how every single day CDPH is leveraging data and we'll very soon come to our architecture, but really the data platform which is on databricks for their day to day operations. So be that as you can see around provider enrollment, right. Monitoring who are the newly enrolled providers to the iis, or reviewing children's vaccination and identifying under vaccinated pockets for children and really taking targeted actions to correct that like school drives for vaccination or reporting back to the cdc. We used to do that initially for Covid, but now we do that for we then graduated to doing it for Covid and mpox and Covid MPOX flu and now we do it for all vaccinations. So one of the other successes through this journey is how we've taken what we built for Covid, a data solution for Covid and scaled that for all vaccinations. Local health jurisdictions are key Partners of CDPH in, you know, championing public health. And our data solution has enabled them to get really unprecedented access to the immunization health and data of the people who live within their county. So again that allows them to do targeted outreaches or actions. And what was in the earlier slide that was represented? There's been almost 8 million residents of California who have received public health outreaches based on the data solution that we have built that could be for COVID vaccinations and urging them to complete the vaccination series. And a lot of focus around vaccination for children under 11 and under 5. Moving on, there was a lot of reporting that we do around equity. Equity is one of the key north stars of the Department of Public Health primarily to ensure that vaccines that are administered, they are administered equitably regardless of somebody's race, ethnicity or demographic profile or socioeconomic privilege. A lot of focus on our data driven A lot of focus of our reporting is on data driven insights to CDPH around equity and then for them to take action to correct that. There is significant work done on vaccine forecasting which is leveraged by the Department of Public Health. And finally, if you go right to the end of the journey, all of this reporting and data is published at multiple public forums. There is a public portal where all the reports are published, there's an open data portal. And of course local health jurisdictions have access to our data platform. And then to end it on like a personal note, we also enabled the digital vaccine record solution. This allows any parent right, such as Jennifer in this example to pull out a digital vaccine record for her four year old. And this has been of immense help for anybody who's ever had to enroll their children for school knows how difficult it is to get the immunization records. This is now available just like in a click of a button and can be sent digitally to their phones. And these are just some of the other examples of reporting that we have done. So moving on to the next piece. So this is the technology piece, this is the shiny piece that we are always interested to talk about. So this is our overall vaccine system architecture. As you can see, it's a very complex ecosystem. To the right, we pull data from several state and federal systems. So federal systems like Tiberius is where we get our vaccine inventory data. And then the main state of California system that we pull data from is Care to which is our immunization registry. And then we bring in other data sets. The Healthy Places index data is primarily for us to measure equity. We also Bring in census data to do different assessment on populations. Our architecture is on an Azure cloud olution. So our ETL standard today is the Azure Data Factory. All of the data sets are put into this gigantic data lake which is on databricks. That's where a lot of the magic happens in in terms of deduplication of data, data normalization, some of the complex patient matching that Michael mentioned earlier as well as some of our more involved AIML solutions for vaccine forecasting. And the Delta lake is really a single source of vaccine data for reporting purposes that really holds all the immunization data within the state of California that is reported by a particular provider into the registry. Now this data set is then further fed into downstream systems. There are several data warehouses and those data warehouses are for specific purposes. Some of them support reporting, some of them support the data for platforms such as the digital vaccine Record solution. Then further on to the right we have again this is a complex economic ecosystem and they all Accenture along with CDPH have worked together to integrate all the pieces so they work seamlessly. But far to the right we have some contact center solutions. Our virtual assistant which is driven by AWS Connect and then the front end is for digital vaccine is built on React js. So again and then yeah of course our reporting is through a combination of Tableau or Power bi. So the slide is available. I won't drain too much on this architecture but really move on to what's our analytics architecture specifically. So as we talked about in the last slide, we have several source systems that we pull data from. ADF is leveraged for doing the ETL and that raw data is loaded into our Delta lake which is really our databricks Delta Lake. That's where we have implemented the medallion architecture really where we take the data through several levels of cleansing and curation. One of the key things that we have been able to innovate through our databricks solution is our ability to do the patient matching. This is a huge problem really like identifying a particular immunization record belongs to a person. X We talked about data quality earlier and if you're thinking of all vaccine data that likely dates back to 20, 25 years prior, we are trying to get like look at all this data and trying to match it to a unique person. So all that matching is happening in our databricks platform and what's done essentially then is these relationships. So the vaccination history of a person is stored in a graphdb. The purpose of that is really quick access to this particular data through an API and then the resulting linkage is again folded back into the data lake. Now the data lake is also sending data to the downstream systems that we talked about earlier. And really our main downstream systems are the data warehouse for reporting. That helps support a lot of the local health jurisdiction reporting as well as the public reporting. We send data to cdc, so that's the other downstream system. And then finally we also send data to the data platform that supports the digital vaccine record solution. So moving on to the next slide. So really our goal here, our transformation journey here, is to move to an analytics as a service architecture. But the reason for us moving there is really anchored first in those basic questions that what did CDPH want out of their analytics solution? It is very simple. They wanted data at really minimal latency. One of the slides earlier established that our solution today allows within an hour of a provider reporting a particular dose into the registry. Our analytics solution makes that available for reporting or consumer access within one hour. That's a little bit, I would say, historically unprecedented because typical latencies was earlier, like days or even weeks. So definitely there was a need for CDPH to have data as current as possible. And then the other piece was for data to be accurate, data to be clean and enriched. And the last piece was the security of the data, that the access of the data was only given to those who were meant to see the data. So really the way, if you see the visual to the right, the way we envision the analytics as a service architecture is really from the Delta Lake. There could be bulk data that is shared with the downstream systems which support reporting, be that public reporting or illegal reporting. But at the same time, we also enable fast data access enabled through APIs, which could either be for just something like querying a person's vaccination record, maybe connecting to a contact tracing solution, or even something like pulling out your own digital vaccine record. If you look to the left, the principles of that architecture essentially is ADF being the ETL tool, replicates data from the source systems into our Delta lake, which is databricks. Any type of bulk data sharing is enabled through the data sharing capabilities. And then we spoke about this earlier. GraphDB essentially used for really fast access to somebody's immunization record. And that's really what's leveraged for connecting to an API. And this really anchors back to, to some of the principles that we had also around the ability to enable fast data sharing, whether it be between the different departments or with public or with other systems. So with this, I will now hand this over to John who will talk a little bit about our short term and long term roadmap. Thank you. Thanks Ajali, really appreciate that. And I'm just going to take a second, even though I know I'm on the clock to appreciate standing up here because I never get to speak and there's probably a good reason for that. I'm not as polished as these guys at all and I use cue cards because you know, I have to have something to keep me on track because of my add. So you know, I'm the closer. Please throw something at me if I get too boring, but these guys have done great work. I'm just here to support them. That's really all it is. And the Department of Public Health's most critical asset in my opinion is data. So we're really working towards building out our data roadmap and what's coming next. So we have to execute on an effective and efficient data delivery model to make all this work. So what you see here. Oops, I'm totally going through my clicker too quick. Nope, we're good. What you see here is the roadmap for the databricks implementation of vaccine management over the next 6 to 12 months which is really comprised of two key phases. The first phase itself scheduled for release in fall of 2023 focuses on building out the vaccine data platform and really for all immunizations from the California Registry into the databricks Delta Lake itself. Some of the capabilities that will be deployed during this release are simple deduplication, inpatient matching, data governance and enhanced network security which is one of my key topics. The second phase of this scheduled for release in spring of 2024 really introduces a robust Master Person Index olution that utilizes AI and ML fuzzy matching powered by databricks Delta Lake and essentially building out a golden record and tracking the immunization history throughout the lifecycle of an individual and most importantly setting this the service oriented architecture framework that introduces data as a service into the entire enterprise itself. So super exciting stuff. While these are the near term 6 to 12 month roadmap, we also have a larger vision as a department and that's currently underway with the future of public Health itself and the IT Data Science and Informatics for a 21st Century Public Health System. God, that's a mouthful. So affectionately known as IT and Data Future Public Health or FOPH for those of you in the state service that always use acronyms. So over the past few years we've really done some amazing things as AJALI and Michael let you know through to respond to a global disruptor which we all know as COVID Pandemic. We stood up systems like Michael said and pipelines aggregated old and new data in days and weeks, not months and years, which is really a huge feat. But now we're really trying to move forward into the strategy side of this. We did it with architecture strategy, so we want to be very prescriptive in how we move forward and invest in our data assets like Databricks itself to meet the health and safety needs of all Californians. Since the work done during the pandemic, we really aligned our efforts to the Department of Public Health's mission and transformational agenda to be able to drive our initiatives including the department wide investment in the future of public health. Like I mentioned, I won't get into the longer name right now and these six foundational pillars that you see before us, but I'm really focusing on the third pillar itself which is the IT Data, Data Science and Informatics pillar, which is a cross cutting pillar and really a key enabler of all of the other five pillars itself. We're taking a strategic approach focusing on the public health key functions and capabilities, also using guiding principles such as IT drives business and business enables IT cloud and Platform first and looking towards modern open architecture approaches to drive some agility and really focus on continuous innovation and continuous delivery for our data itself. The last thing I want to talk to you about is that big long name I said or the future of public health and the nine initiatives. The nine initiatives itself really align to two key objectives. The first objective, and I could read it for you, but really what it is is it's IT and program working together to really be more rapid and advance the public health goals itself. And again we've done this in silos throughout the past. We're really trying to bridge that gap and use IT and business to work work together. The second objective is create a valuable agile cost effective again won't read IT technology ecosystem but what that really means is make prescriptive investments in what we need to be able to be responsive to the next global disruptor itself. These initiatives are really important to provide timely, accurate data to make effective data driven decisions so again we can be responsive to what the future holds for public health events. Additionally, these initiatives will support the creation of a modern cloud based agile and responsive infrastructure, identify and prioritize critical data and IT systems to respond and scale to public health events and create a cost effective technology foundation. That and these are important words, standardized reusable and leverageable for enterprise services and enterprise capabilities. We have really stood up, like I said, and responded to this global disruptor by creating these systems. But now it's time to really look at our total cost of ownership and our technical debt that we've incurred to be able to respond to this and be responsible about consolidating those and look at standardization, leverage and reuse itself. So to ensure that the department's successful going forward in this adventure, we have done some massive change within it. And one is create a IT operating model that is not siloed and that is customer centric, including an enterprise data analytics branch. We've invested in upscaling our workforces so we can have the knowledge, skills and ability to be able to support and maintain these new capabilities and environments that we're standing up and finally really create a robust data governance framework so we can make informed data decisions, prioritize workloads, address risk and evaluate business value realization. So we intend to continue moving forward on this data journey, looking for ways to be creative about optimizing our data, about maturing and transforming our ecosystems to really meet our goals for the future of public health itself. I am committed to this and our entire department is committed to this. And that is one reason I love working at the Department of Public Health is because we all have a collective goal and I feel like everybody is there for the right reasons and really that's to work for you as Californians or if you're not in California, look at what we do for Californians. So thank you very much. Like I said, hope I wasn't boring. I appreciate it. And that's it. Thank you. Are we ready? Bill, Amazing job. Amazing job. Given we're a little short on time, I'm going to get us out to break. But Michael, Anjali, John, thank you so much for your innovation. Thank you for sharing your insight. We have an amazing presentation when you come back. Veteran affairs within 15 minutes is going to be the first. Then we have a federal panel and the recognition for our recipient of the Public Sector Data team. So you want to come right back and then happy hour. We'll celebrate all your success. We've got an exciting conversation with our friends of Veteran Affairs. I will let the good folks at VA introduce himself but real excited about this conversation. I think you're going to really enjoy it. Brian, why don't you introduce the team. Thank you everyone for coming back. It's always a risk about having a break 2/3 of the way through the program, so thank you and the good news is we now have, I think, room for everyone who wants to sit. It was a little crowded in here before. Bad news is for everybody who left, they're going to miss the best panel of the day. Playing to the audience right here in the front. I'll pay you later. All right. So yeah, as Jude said, we're going to be talking about Department of Veterans Affairs. Given the size, the scale, the breadth of use cases and the fact that VA was an early adopter of the cloud, you guys are a leader in the Fed space whether you know it or not. And as Jude mentioned earlier, right. VAFSC DAS is a finalist for the Data Team Disruptor Award. So that award will be at 6. So hopefully we'll grab a drink, head downstairs, see you guys on stage again and then, you know, over the course of the panel I hope to highlight a few themes. Fed Health, Data sharing, Analytics and Generative it. So to introduce myself, I'm Brian Davis. I'm Director of Federal Civilian Civilian Sales here at Databricks. My team focuses primarily on the healthcare accounts. I'll turn it over to you, Tom, to introduce yourself. Super. My name is Tom Latimer and I report to my service director Scott Meyer, who's sitting out here for the Data Analytics Service which is part of the Financial Services center at the fsc. Great, thank you, Joe. Hi, I'm Joe Daria, Director of Data and analytics in the Product Engineering Service. I work for the Deputy CIO of Product Engineering. She Right over there, Ms. Carrie Lee. I thought that was better than just reading your bio, so appreciate that, thank you. Fair enough. All right, let's get started. So Joe, can you give us some background on your program in the sense of, as I mentioned before, the scale and the scope of the program?
Absolutely. So we're blessed at VA with what I think is probably the best mission in the federal government and that is taking care of veterans. And what we do in the Data and Analytics group in, in Product engineering is look at the full scope and totality of what that means. So we look at data from DoD that describes what a service member did, where they were, what jobs they performed to, ensuring that addresses are synchronized across various veteran facing and employee facing applications and a number of things in between that facilitate real time data sharing. Right. But what we're really focusing on is the modernization of everything that we've done. Because, because as you mentioned, we have been a leader in a number of areas and self service data is one of them. I mean for nearly 20 years the corporate Data Warehouse has been aggregating a holistic picture of health across the Vista systems in the VHA community and 130 odd hospitals and outpatient clinics. And we're taking that from self service analytics on prem, as you said, to cloud scale and we're doing that inside the system. Summit Data Platform and the Summit Data Platform is combining a number of work streams that we had started in various functional areas that support measurement, analysis and everything in between for veteran care and service. Our CX Insights Group and the Customer Experience Data Warehouse hdap the health data and analytics platform and pulling that into one cohesive unit that has standardized patterns, access controls and general rules of the road that your average data analyst, data scientist or just plain executives that need to see power bi to understand and measure the quality of service delivery in their organization in one place. Great, thank you Ed. Tom so our journey began back in 2015 as a kind of a proof of concept and had some wonderful early successes and became an actual service and 2017 and started out with a staff of about 14 combined federal and contractors. And we've grown that business very successfully over the past several years to about 30 federal folks and 60 contractors. So done very well with that. Our scope is that we've got one of our biggest benefits is our access to data resources and things of that nature that we can parlay together to transition data into usable decision informing information. So that's kind of where we're going. And we've transitioned from using ON PREM in the process of transitioning from using ON PREM to more cloud based things and some hybrid capabilities as well. Great, thank you. I think now that we have a sense of the scale and scope of your projects, maybe we can get down and get an idea of some of the use cases that your teams are enabling. I think that's probably the thing that most people in the audience want to hear. So we'll start with you Tom. One of the great things that we did, and I should have mentioned this a minute ago, is three years ago we initiated and Dave Fuller is the author of this, the guy sitting in the second row here, a hackathon. We call it a data thon now, but at the time it was called a hackathon. And we did had two very successful years. Last year we had just a real neat use case where we identified components that people were trying to purchase to support medical operations. And one of the challenges is there's a lack of uniformity in how they approach that, how they approach that purchase. There's a product and there's a contract called the Medical Surgical prime vendor list, and we like to see greater use of that. There's a lot of efficiencies that come with that. Not just financial, but just operating efficiencies and things like that, being able to handle crises like Covid or be more effective in crisis like Covid. So at the onset of the data thon, we identified that there was only about 25, 26% use of the Medical Surgical prime vendor list. And the data thon identified that, in fact. In fact, we could take purchases that are taking place outside the MSPV and align them with functional equivalents that are on the mspv. We moved that forward into a rapid prototyping, and we were successful in adjusting the match of those component items from about 26% to just north of 60%. So that was a very successful program for us. A couple of other things. We did some work with the National Artificial Intelligence Institute, kind of two use cases there. One, we used veteran data to identify the alignment of comorbidities with different types of treatments and how those treatments were successful over time. And then there was another one where we used a tremendous amount of data from congressional reports over about 22 years, using natural language processing to identify similarities in those reports and to begin to forecast what future needs, you know, data analytics needs, might be requested of us. So those are some of the use cases that we've deployed and worked with, right? Yeah. And I was actually involved in the initial hackathon we did two years ago. And you have another one coming up in October. I just wonder, just for the audience. Right. Do you feel like that's a good tool to kind of raise awareness and enablement of the platform? Obviously, you guys get great outputs, but from a user awareness standpoint, do you think it's a good tool? Thanks for that question. I think that's an amazing tool. Again, this is our third annual hackathon. This is our third effort in doing this. And what we've been very successful in doing is bringing together people from across industry, from across the government, not just within the va, but the really smart people who can solve some of these problems. And so we, you know, we're in the process of setting up for it now. We've not identified a use case, but I hope we can find something that's going to be a real challenging thing for us to tackle. And it's very exciting to do that because it's not just the breadth of the people that come and join us with this, but the resources they bring to the table their experience and their capabilities. They bring to the table, they teach us and bring us to another level. And also, heck, you know, we learn a lot from them and, you know, it changes how we address and approach different things. This last year was a great example as we moved from the data thon to the rapid prototyping and now we're operationalizing that concept for them for the vha. That's great. Yeah. As I said, I was involved in the first one. My team won. I retired on a high note. Maybe I'll come out of retirement. So I was going to say, are you going to be back this year? I think so. I'm excited about it. Have you registered is the question. I got the invite. Okay. Well, there you go. All right. So, Joe, can you talk about some of the use cases your team's enabling? Yeah, I think we've got two general themes that run through many of the use cases that we bring to summit. There's first, think about CDW, the 10,000 SQL writers, the 100,000 data consumers across VA, massive scale on an on prem system that's mostly focused on primarily descriptive analytics, power BI reporting, pyramid analytics, types of things that kind of tell you the state of the world. And as they migrate towards cloud native solutions, a couple of use cases of note that kind of take that into the next generation of where we want to be in terms of being out in front of veteran needs and being proactive about taking care of veterans first. One that comes to mind would be the long Covid use case that we've been working on on that one of our Presidential Innovation Fellows has been driving, which is looking at using NLP and other advanced techniques to identify veterans that may be in need of long Covid care and proactively reaching out to them and saying, we have a belief that you may need additional screening, you may need additional care and support. Is that true? Kind of. In that same vein, the Coordinated Care Tracking system, which was born first out of cancer care tracking, looks at computable health data and unstructured health data and evaluates veterans that may be missing care in terms of things that should have been scheduled that were not veterans who have missed appointments or veterans who may benefit from a course of care that is more aligned to the stage of their particular malady along with comorbidities. Right. So looking at the cohorts of patients that are in that same area and then conducting that outreach, we've done some great work on accelerating some of the metrics that we use in mental health and suicide prevention looking at things as simple as medication fill rates. In other words, are people compliant with the prescriptions that they're on? We're always looking for those early indicators of things that is a silent indication that there's something wrong with this individual that care and intervention can help. And then if you even look back to the things that we're trying to accelerate from on prem with our reach vet applications in storm looking at missed appointments again looking at those patterns of behavior that indicates somebody may be in crisis, there's a direct measurable outcome of 5% reductions in suicides just from following up on people who missed their appointments with vha. So I think when you look at the totality of the data that's available on veterans across our benefits administration, our healthcare delivery, there is tremendous opportunity for us to continue to look for these patterns and indicators that are going to save lives, improve veteran service and allow us to be more proactive about targeting benefits and targeting care to those that may not even know they're eligible. And to tie it back to the day in AI Summit, I believe believe you guys won a Data for Good award for that program. I did hear that but you know, I didn't want to brag and secondly, you know, I was already a little upset that there was no fire. So you know, virtual fire on the screen. So maybe I'll tee Joe up with this question. So Joe's actually speaking on Thursday if you want to hear more and maybe he'll give us a little preview you into his session at 2:30 on Thursday on Lessons Learned from VA Summit Data Platform journey. So the question is what are some of the biggest challenges you've faced so far? So getting back to I think Arsalan's initial kickoff about you guys are the leaders, you have paved the way. You've overcome some of these challenges like help some of our friends that maybe aren't as far down that journey. I think there's a couple things that resonated with with me in that talk and listening to others in some of the other presentations today really comes down to your people, right? If you're aligning your people up for success, if you're thinking about upskilling them, if you're clear on what your strategy is, if you're giving them a North Star to align around, it doesn't matter if it's data, it doesn't matter if it's baseball, right? You're going to be more successful if everybody understands the rules of the game. That's been particularly I have a particularly large federal workforce that's a little bit abnormal in federal IT of folks that built CDW and have taken care of it for years. And we're at the forefront of data management and analytics development in the federal space. And now I'm taking those same employees and saying, okay, this was great, it was amazing, it was a breakthrough. Nobody else was doing it at this scale. Now we're going to do it again. And that can drive a little bit of scale skepticism because it runs va. If CDW were to disappear tomorrow, can't run va. And I tell people that all the time when you talk about the importance of some of our legacy systems and the impact that they have. But at the same time, there's a lot of patterns and ways of working that exist around that platform and that ecosystem that are not optimal. There's a lot of copies of data, there's a lot of, I would say, ambiguity in terms of how can I get to a new tool set around the state. And the answer is, you really can't. On prem, there's a lot of things you can do that replicate some of the things that you can find that we have up in our lakehouse architected platform summit, but it's just not the same. And so building the bridge between where we are today and why VA must evolve because just the great growing amount of data that we collect about veterans, everything from patient generated health data to health data monitors to imaging, just think about how all of that continues to increase and with the resolution and improvement in all of those devices at the point of care. Right. You know, we have to look for those opportunities to bring those insights to clinicians, bring those insights to people who are managing supply chains, bring those insights to people who have to talk to Congress and, and explain why VA needs the funds it needs to deliver world class care to veterans. And all of those challenges then are wrapped up in the fact that I'm still in a very large, over 400,000 employee federal agency that has a very particular way about doing everything, whether it's provisioning things in the cloud, whether it's ATO management or whether it's procurement. And all of those government realities I deal with while we're trying to push forward with fundamentally changing the way a large group of federal employees work and bringing another 300 contractors along with them. So I would say that, you know, not to hit this too hard, but you've heard it from everybody. It's people process technology. The technology is the easiest, even in government. So that's, I Think our lesson is you've got to have a, you got to have alignment amongst the team and you got to be able to bring, bring the knowledge, skills and upskilling and reskilling to them or you can't expect them to do it on their own. I mean they've got a day job. So you've mentioned CEW several times here. I'm not sure. People outside the world. Sorry, government guy speaks in acronyms the corporate data warehouse horribly named Solution. It was really vha's data warehouse that explains expanded into some other roles. And VHA is the Veteran Health Administration. We also have the Veteran Benefit Administration and we have the National Cemetery Administration and then about 30 other staff offices with incomprehensible names. So what I was going to ask about the cw. I've heard that possibly it's the largest data warehouse in the world. Not here to dispute that. But speaking of size and scale, can you give us some idea of what that looks like? Like, I mean you're talking about multi petabytes of data and if you think about the fact that you're, you're dealing with patient encounters, right? We talked about how, you know, there's, there's, you may have heard, you know, that we have this EHR system out in the field that VA developed called Vista. And Vista was at all of our medical centers and outpatient clinics and it's configured at the field level, at the VISN level. So the joke is if you've seen one Vista, you've seen one Vista. So now if you take the Vistas and you want to build an enterprise view of what's going on in healthcare, you have to rationalize stop codes and coding everything else. Now, I'm not saying that an ICD10 code in one place isn't the same as an ICD10 code somewhere else, but all of the business rules that may happen within that hospital can be very different. And so then when you bring that up at national scale, you have to rationalize all of that. So if you think about, I bring raw data in, I transform that data here, somebody else models it into this, somebody else has optimized that same data set for columnar high speed analytics and reporting. You're going to have a couple copies of the data somewhere along the line. And that's where some of the size comes from. Then some of the other size comes from the fact that right now there's probably 9.2 million patients receiving care in VHA. Dr. Scott might dispute that. He kind of shook his head a Little bit. It's probably higher. Probably higher, but close enough. But remember, so the history for every patient we've ever seen in the Vista era is also in cdw, whether they're an active patient or not. And then that supports the research community, so on and so forth. All right, thank you, Joe. Tom, what are the subjects of the biggest challenges you guys have faced? So first, I want to echo something that you brought up and, you know, the people process technology discussion. You know, for us, the first thing is all about people. And your topic about, you know, building the right azimuth for those folks to follow to empower their success is huge. You know, and that's. That's been one of the most gratifying things in the past five years that we've been able to build out within das. The process piece of it. I mean, how the difference between one station or one hospital versus another handles the use of that data and the collection, the use and application of that data and how they try to bring about better outcomes, that it is different is challenging. It's tough, and it's tough to deal with. But I would like to speak to another piece of it, and that is you're asking about challenges. And we really have. One of the things that I think is, if I can turn this question around, a huge benefit. And that is the ability to operate with a great deal of agility because we're a franchise fund. So we've got customers who can come to us in fairly short order without going through an appropriations track or going through their own contracting track. We can turn around and do work for them fairly quickly and deliver good outcomes and good informing information or good use of their data for them to help them inform their decision. Now, the challenge with that is a couple of things. One, people don't understand what franchise fund is, right? They don't understand it's a fee for service type of a product. So that's a bit of a challenge because it's not the normal appropriated type of work that we do. The other is organizational understanding, you know, not just the people that we work with, but just the organizations in general and how they, you know, how they view franchise fund because it's different and things like that. So those are some big challenges for us. Understanding of the data. That's kind of one of our mantras here. You know, data. I mean, if you can think of data as this big jumbled mess that is very, very difficult to understand and very difficult to make useful to speak to the question or the problem, that's Trying to be answered the that is one of the key components. And being able to speak to that, being able to understand that data, being able to apply it to their business need, most importantly, understanding that business need, what's the problem they're trying to solve, how do we tackle that, how do we resource the right and most economic resources to tackle it and things like that. Those are some of the challenges that we face. But you know, we're moving forward. I think the future is really, really bright. There's a lot of great technological advances and even more importantly, the folks that support us and the folks that are on our team back to people are building every day their skills to be much more effective in how we approach these challenges. And Tom mentioned that maybe not everybody's familiar with the VA franchise fund. He has too much humility, but I don't. So there's a brochure down here for anybody who wants to learn more about the va. I'm glad you did that. I completely forgot. Oh, thanks. I appreciate it. I just flat brain farted that. Sorry. So given the theme of the conference around generative AI and AI in general, maybe you've heard a little bit about that in the news the past two months. How do you think the VA has been able to take advantage of the latest developments in generative AI and AI in general? And I think even as long ago as last year, VA did some presentations on LLM. So I mean I think you guys are at the forefront of this, so I'd love to hear your thoughts. Yeah, I gotta go back to the outcomes of our last data thon and how we moved that from, you know, proving that we could do those functional equivalents to being able to do a rapid prototyping where we identified not only can we make those matches happen, but we can also continually improve the rate of those matches. So it's an ever improving system that is going to move us more and more towards that MSPV adoption. And frankly it's, you know, our hopes are that it's going to inform the MSPV in future states so that we have more useful line items on the MSPV that are applied to the direction requirements that we have and things like that. So, you know, and if you think about it, it's vha, it's big, it's big, big, big hospital system, things like that. We buy a lot of stuff every year but the challenge with that is making sure that we buy the right stuff and we set conditions for the success in all those facilities so that they can address oncoming problems. They can address, you know, the next COVID 19 and things like that. So I think the future is very, very bright for this. Obviously, I think it's got to be responsibly managed, but just these few examples that we've had here, I think are going to propel us forward. Well, and I think it's, you know, we've got to move very quickly away from just being the descriptive analytics that he talked about to being much more predictive and get out in front and, you know, understand what the future holds and be in front of things so that, you know, our practitioners are prepared, they're ready to deal with patients and crisis and help veterans be more healthy. Yeah. I don't know if you've seen the slide. We have the data maturity slide. Right. And the first three steps are kind of tell me what happened. Right. Looking forward, the predictive and the AI capabilities where we need to be very much so. So, Joe, around generative AI and AI in general. Yeah. So I think our cto, Charles Worthington, is definitely got some of his best people fanning out across VA right now, delivering the message around how we become more ready to do that at scale. Right. Even in my area in and around Summit, we've got a couple use cases that are taking advantage of that to kind of push forward in a trial basis. My big concern is, and this goes back to kind of that maturity model piece, how confident are we in the quality of data that we're starting to ingest from a multiple of enterprise sources? We're in a heavy modernization phase right now at va. So that means a lot of systems that were Vista based, not just healthcare, but, you know, logistics, other management systems, financial management are all in transition. And when that happens, we have to make sure that we've got the best view of that from across the enterprise and that we've appropriately cataloged it. I think the Advanta folks talked a little bit about the Federated Data Catalog. My team executes that on the VA side in partnership with architecture and engineering services. And that's foundational to us to make sure that we are understanding the metadata behind everything that we're processing on these systems, that we're adequately providing transparency to potential users. I made a really bad joke, I think, two briefings ago to the cio and I said, well, you know, the great thing about AI is it can allow us to make really bad decisions faster if we're not careful. Right. And he laughed after he got past it. But the reality of it is We've been very deliberate and intentional about how we curate and make data sets available for enterprise consumption. And we need to make sure we maintain that disciplinary. As we move into more of the AI world, we've got a lot of folks who are pushing hard on it internally. Smart folks doing great things. I think one of them's here. I haven't seen him yet, but Schaefer. Oh, there's Schaefer. All right. So he's doing that in and around our platform. So I think we're at that. I think we're somewhere between crawling and walking. And what we want to do is learn a lot from that, build a little, test a lot to ensure that as we move out from there, when we get to that run phase, we're doing it in an equitable and sustainable manner. Okay, great. Thank you. So we'll try to end on a high note. I asked earlier about challenges that can be sometimes construed as a negative. Right. But we'll end on some best practices that you can impart to our friends here. So closing thoughts or top recommendations for other large organizations trying to modernize. Start with you, Joe. Yeah, I would say I'm going to. Going to tread some of the ground that we've been on. Right. And that is clear strategy, an intentionality of knowing where your people are today, where you want them to be, making sure that you're giving them the resources to do that and get there. Whether that's training, whether that's coaching, whether that's finding the right contractors to help them, that's key. Supporting them through that journey and understanding that they have a job that you're asking them to essentially throw away maybe 10 or 15 years of experience and do something differently. I spend a lot of time in the commercial world. I always joke that I had to relearn how to do my job every three years. That's a culture that we have to instill in the federal workforce. We're very blessed at VA that our top leadership in OIT is very much in tune with that, that we need to find the best people possible, get them in the door. I believe there may be a little event in Palo Alto on Thursday that's focused on that to help us find some more smart folks to come in here and help again. Beyond that, beyond that training, I think that the issue really comes down to is you can't have one of everything. You're going to have to pick some winners and losers in the technology space. You cannot integrate every tool available just because somebody jumps up and down, down about it. And they saw a really cool social media ad and think, oh, VA should be doing that in the enterprise. Done work like that. And you gotta be disciplined to stay in your lane and incrementally build out those capabilities and know that the first part of your journey isn't going to be incredibly flashy. Right. It's gonna be focused on instilling those data management practices, being disciplined about how you get to a data catalog and building repeatable patterns that are going to take teams that are just on their early journey to doing cloud scale analytics, that they see that there's a paved path and a way to get there. Great. Thank you, Tom. So two components and I see we're running out of time, so I'll try and be quick. Two components. First, I want to echo what you're talking about as far as the people and discipline aspect of this. Very much so. The not always jumping to the newest flashy toy that's available. And so often we see it on social media, but we jump to those types of things oftentimes without vetting them to see are they really providing greater value or are they just new? Do they have a better GUI or is it just new? What's the difference? So that's one component of it. The second would be VA is a huge organization for 450, 460 odd thousand people, second largest organization in the United States beside the Department of Defense. One challenge that we have is all the disparate efforts to try and bring about positive outcomes for the veterans, all the disparate efforts to try and bring about to utilize analytics to do exactly that, to have better outcomes for the veterans. And I think one place that we can improve and work a lot better is by, is by working together, not having disparate things that are just out there that don't comply. They're not covered by governance, those types of things. We need the governance to be an enabler, not an inhibitor. And that's how we need to look at it. But also where we do have like efforts proceeding. How do we collaborate and complement each other and work together to the strengths of both organizations to do better? Anybody can build a tech stack. We've seen tons of them pop up in the past couple of years. And one of the things that I would close with is, you know, it's not the tech stack, the tech stacks, neat, it's cool, but it's the people and their skills and how well you've weathered those lessons learned and how well you've worked yourself through Those lessons learned, that's going to define success. Great. And we're right at time. So I want to thank Joe and Tom for your participation. Hopefully you guys will be at the happy hour afterwards to help answer questions. So thank you. Don't forget the flyers if you need some help. Thank you. Thanks, Tom. Thank you. Nice job, Brian. We have a real treat, our last panel of the evening, the federal panel. Kristin Nero, can you come up with our friends from CIS and from Coast Guard and from U.S. postal Service Service, why don't you welcome Dan Houston, Captain Brian Erickson. And from CIS, we've got someone who's ably standing in. Yes, yes, yes. Thank you so much. And then after this panel, we have our pub Sec data for good recognition before closing out with happy hour. You're going to want to hear that. Thanks, Jude. So lovely to be here with you all today. I think since we are standing between happy hour and maybe we'll go ahead and kind of kick it off. I'll tee up the first question and you guys can introduce yourself as a part of the response. That would be great. So, you know, we heard a lot about technology is not the problem. Right. And I'll let you guys decide what the problem is. But I know you all are in different parts of your journey in modernization, so if you could just kind of, you know, tee up for the audience. Your where you at are at in your cloud migration journey today, in your modernization efforts today, and really what you, whether in your design or in what you're seeing now and you're trying to like, maintain, what are the principles to that effort? Dan, I'll start with you.
Sure. So I'm Dan Houston. I manage the data science and exploration team at the Postal Service. We're a small group. We're largely looking for patterns and data that would not otherwise be discovered. So we're a huge organization, lots of analysts spread everywhere and they can find the things in their lines of business, but they may not be able to find the things across lines of business. So that's where we step in and help. We also run a large data lake. So that was sort of the beginning of our journey. So getting tools, technology, data to get in one place for people to start answering the questions, because that often took months of them trying to do before they could get started was our first step. Our next step is where we are now, which is looking at getting into the cloud. So how do we do it faster now that we have data together, now that we have tools, technology and people working together and Getting same or similar answers. How do we do that faster or be able to run more experiments? And that's where cloud elasticity is huge. Right? Cloud for government usually is not a cost savings. That's how they try to sell it to us, but it's usually not right. So what is the other opportunity? It's that I can spin up a lot more fast. I can now say, do you want me to do it in two hours or run 10,000 versions of that? We can do that. And here's the cost associated with it versus I can't say that now. I can just say, well, we got to wait till the end of the week to get you the answer. So that elasticity is huge and then bringing on new technologies, right? So everybody has really covered sort of the people process technology thing really, really well. I think for us, when we're looking at the platform, we're looking at something that's simple, something that's smart and something that's reliable. And so by simple, you know, when you look at these technology diagrams, they're not simple, right? But keeping it as simple as you can, keeping it something that isn't going to allow a vendor to lock you in because it's so complicated you can't possibly undo it anymore, keeping it smart and that to me, making sure everybody's talked about the shiny object, staying away from that, not letting it distract you for sure. And also the vendors that you're partnering with, who are you getting married to and how long is it going to last? Right? So finding those good partners is to me more important than the technology. And last is reliability. Technology gets blamed for unreliable data and an unreliable platform. So you will ruin a relationship really, really fast if you're not reliable, if your data's not good, it's not going to be great, right? We're never going to have great data. It's got to be good and it's got to be available. People got to be able to use it. So that's it for the postal Service, I think. Thank you. Hey, how are you doing? Hey. Captain Brian Erickson. I'm the chief data and Artificial Intelligence Officer of the United States Coast Guard and I would say that we're kind of early on our journey of digital transformation or really kind of data transformation. So we like to say that we're the longest continuous seagoing service in America because the Navy kind of broke up for a little while. But so 233 year old organization, you know, there's some data problems in there over the last 40 years. You know, it was kind of the information era, the rise of the cio. And then only about three years ago did we recognize that, hey, this is not just an IT function. Data's more than that. It's kind of a core business function, and each of these business verticals needs some sort of a support. So we started off with commissioning a data readiness task force for which I led, and then ultimately that resulted in making me as the first Chief Data Officer of the company Coast Guard only about a year and a half ago, and now just trying to kind of keep pace and also signal to our workforce that we're paying attention to this. I changed the or we changed the name to the Chief Data and Artificial Intelligence Officer just only a few months ago. So, you know, our journey in data and artificial intelligence is very young, and we're behind. We rely on a lot of partners like Nick and Cody. And the team from CDAO are kind of my peers over in the. The DOD camp, similar peers in the DHS camp. I feel like a lot of times I kind of have my foot in two camps, which is pretty common for us as an armed force, the only armed force in dhs. But it's for me in my time, and in about 10 days, I go on my 31st year of active duty service in the US Coast Guard. For me and all the jobs that I've had, this one is the one. Thank you, thank you, thank you, thank you, thank you. But this job in particular, I feel like I work kind of in the DoD space. We are going to support the warfighter. We're going to be able to fight with the warfighter, and then also in the DHS space as well. So that's kind of where we're at in our journey. Thank you, Prabha. Good afternoon, everyone. I'm not Sean Benjamin. I'm just representing him here. He could not make it today because of the east coast airline issues. So I'm just here. I just got to know two hours before that I have to speak. So I'll try to do the justice as he does. I've been with USCIS for almost 10 years, and I closely work with Sean. I'm with a program called Data and Business Intelligence Section for uscis. I manage the Enterprise Data Lake for uscis. So if. If I have to start. How did we start the journey of modernization? If I have to say that, as Arsalan mentioned, uscis, they started with a poc. I was one of the member who even started that groundwork. So six years ago, we Were only three member and one of my colleague is also here. We were all put into a room and we were called as Tiger team and they asked us just do data breaks. And like anyone else, change is always hard, hard. Which was hard for me because I'm a legacy person, a data warehouse person where I like to stick onto my comfort zone. And I said everything is cruising fine. Why are they even asking us to change? But they just said, let's try this. We are already in AWS cloud. Let's go and try out databricks and do modernization. Okay, what is this modernization? Six years ago it was a new jargon everyone was using. So that's how the journey started. And exactly are our data warehouses mostly ingestion of data from disparate systems, load it into your Oracle targets and push it to the downstream. That's a simple logic, that's all we had. So they wanted us to see how can we remove the pain points which we had that day. Like can the users have the data availability and the data usability can be faster, can you have more relevant data or a real time data? It was all new words put into our room and they said we're going to try it today. And all three of us we sat and we are more a Jira GUI person. We cannot code that extensively in Scala they said Scala, Python, that always a new term for us. But we just started with the basic extraction, load and extraction and loading of a team. When we saw that was only three lines of code which can do everything and it was matter of hours, which in our ETL processes who do it in hours can do it in a minute. That was the magic. That's where we all said, oh my God, we can do so much in five minutes. Then why not the whole legacy system can be changed into the modern platform. So that's how we started this journey. And the main pain point for us next was, okay, we know this works. How are we even going to give or educate this or bring our team up to speed? People all wanted, no, we cannot learn Scala. That was another complaint the team was mentioning. So we took small teams and we started educating them and brainstorming them and slowly we gave them a parallel pipeline where they can onboard whatever they are experimenting with this new Scala or the new databricks platform. Load all the data and then you still have your Oracle footprint and you can still compare and you feel that comfort level. Right. We kind of gave that opportunity for them to transition and then slowly today we brought in new applications like Tableau and we siloed the users in our enterprise data lake hub house to different part. Like the sophisticated users can be data scientists or even a basic BA can use the databricks or a tableau as an application. And we kind of catered. So that's when we saw the journey which we started six years ago was matter of thousand or 1,500 users. But today we have 15,000 users altogether in the downstream applications. And we can work seamlessly and we can reuse many of the pipelines, whatever we have built and we are more adaptable to change. And yesterday again we saw the LLMs. That's what I was telling my team as well. It's going to make us more lazy. That's how LLMs are. Right. Everything is there. So again, that's another change which we all are overwhelmed already and we are still adapting and we are thinking how are we going to onboard the journey more lazy, more productive, so getting more for less. Right, well, and you just touched on it and we heard it, you know, with the Advanta team today. Right. I think the key to a lot of these large modernization efforts is being able to meet users where they're at. Right. And so, you know, we pride ourselves, obviously we put open source above everything right here at Databricks in order to ensure that like you, you know, you can literally let people speak their own language. Right. So how have you seen, you know, the growth of open source affect and provide, you know, the ability for, you know, more productivity or more value or more strategic insight in your agencies? Sure. So at the Postal Service we definitely leverage open source technology. I think the biggest advantage though really has been the pressure it's putting not only on other open source vendors, but also on commercial vendors. From my experience, you started to see a lot of stagnation. And how do I just keep this customer and not really how do I innovate? How do I move the ball further to make my customer more successful? So open source brings that pressure. Right now anybody can innovate without buying anything. Right. I can just start, I can just play with it, I can experiment with it. The Postal Service has embraced some development areas where we allow that kind of software to just be brought into the environment, keep it away from our production spaces so it's still protected, but allows for that innovation and that's huge. Right. Once we do that innovation, you've got a better use case, a better business case to bring forward and actually seek a solution in our environment. So for me, while it's huge in the innovation Space. It's huge in just being able. Anybody can, can pick something up and work a problem. You don't have to go buy software to do it. It's that innovative pressure that's so key to me. Yeah, like failing forward without the financial risk. Absolutely, absolutely. And back to your ephemeral comment. Right. It's pretty, it's exciting. Captain Any. So I would say that's the future that we're driving towards is. You know, I haven't been burned by, by that issue yet but the idea of being able to just to have different business verticals, whether it's finance or operations or intel or cybersecurity logistics, be able to just create with tools that are brought to him within an open source type of architecture is what we are trying to create right now. Awesome. From our point of view, Open source is the bigger one was the Delta Share concept. So the biggest pain point after we modernized was how can we share data between inter agency. So that was the bigger point. One is the other agency who's going to receive this data may not have databricks in their agencies or they were still waiting or relying on the old way of transmitting data through files or through emails, which is old technology. We implemented the Delta Share with the help of databricks team and that's one thing which opened the gates for us to not share the data physically to the other agency but at least make a pipe or open it to them. With all our governance and security in our premise, whatever we have kept our data the secure way. However we have just give them only the select access. So we recently did it with CBP and only the configuring of that Delta Share concept was. It took some time but we are at a point where it can be reused and any number of any tables or any data can be shared seamlessly. So. So that's one concept. And also we are exploring on the concept of Open share where when the other agency is not having Delta Share, it's not having a databricks, we can still share with them through a concept of open share. So that's something which we're exploring. Yeah, I think the Delta sharing piece is just data sharing, secure data sharing in general. Right. I mean that's what a lot of agencies need to do. You need to. To do it with your partner agencies with even maybe partner private companies. Right. There's a lot of value to it. But being able to provide access as opposed to like you know, sending a copy of data and then solving the for the data staleness and Data integrity issues and also auditability. Right. It's exciting. So, you know, one value of open source that I think I would imagine would matter largely to the Coast Guard and are in our, you know, military affiliates would probably be, you know, the value to leverage unstructured data. Right. How do you now that you can apply analytics to more than just structured data? Right. That was kind of like in the traditional data warehouse world. What do you see could be the power of that? Captain? Yeah. So whether it's unstructured or structured data, I think the power that we're seeking is speed to create. Like we're a response organization. When a disaster happens, we surge and we often search quick and sometimes it will be in a scenario that we haven't encountered before. Let's take hunting for submarines, searching for submarines up in the Northeast. There was a number of folks that had reached out to me to figure out if we could use some sort of AI activity or analytic to kind of determine maybe a little bit better in the water column where you might find the drift patterns occurring. And we don't typically search in vertical columns. Of course, we have partners that are prepared for that, and so we would partner with them. But within the organization, I think that what we need to create, whether it's structured or unstructured, is this foundation that we have yet to create. And we talked about the people, processes. I also kind of roll culture into that and technology, that whole foundation of the people and the scientists that are ready and prepared to react to some sort of a contingency activity or operation that we're working on, or even a more business data mining effort with speed. We gotta have that foundation, that technology foundation, the foundation of people, the processes in place. And that's kind of what we're working on. Awesome. So not to be cheesy, because I feel like I don't know how many times generative AI is going to be said during this conference, but maybe we could get a ticker going. But you know, where I think the value is, obviously to each agency, I'm sure, is somewhat specific and also endless. Right. And where does someone start? Right. How do you begin the process? I think it was Mr. Bang that said start small, which makes a lot of sense. Right. And then maybe you can build those principles. But, like, where do you intend to spend, spend your time and how do you get started?
Dan? Yeah, sure. So at the postal service, we have started with generative AI. Certainly we're playing with some of the GPT models. You know, I think for me, the big Thing is, we're going to be forced into this space of having to use this technology without being experts at it, right. So we're going to have to start using it as organization. People are going to have that expectation. I joked with somebody earlier today, I don't remember which test it took, but I got a 70% on it and everybody was excited about that. If I deployed anything that got a 70%, people would shoot me. So finding where those edges are, where you're going to fall off the edge, where these models answer very authoritatively, it is this, it is that. They don't say, I think so. Humans read that and to the laziness comment before it's true, right. I mean, we believe it. So no matter what questions you ask, if it says it's this, we will tend to believe that. And if we make dangerous statements with that, that's awful. If we just make wrong statements, it leads to distrust and ultimately destroys our ability to leverage the technology to its greatest extent. So for me, it really is finding those edges. So for a long time I spent trying to make them say I don't know. They have prompts that you can generate to say, if I don't tell you enough information, say I don't know. I had to take one all the way to do, you know, Dan Houston's son, to get it to finally say I don't know. Even for me, it said it knew me and made up some answer. Right. So those are called hallucinations in these models and they will do them very authoritatively, right? So touching those edges and knowing where you're going to fall off, knowing how to make them safe, say I don't know. Don't just trust because it gave you a right answer of something you fed into it, that it's doing a good job. Make it say I don't know. If it won't ever say I don't know, you should not deploy it. Right? Because it's going to give people bad information, it's going to give them bad answers. So for me it's that, right? Obviously we got to get better at it. We got to get better at the expertise. We got to get people who know how to do prompt engineering, we got to get people that know how to fine tune. We got to get people that understand these vector stores better. But in the short term we're going to have to use this. So figuring out those edges for your use cases I think is critical. Yeah. And I think we naturally. I like that comment about the edges. I think as humans we naturally do that. Right. There was a day where there were doctors that didn't use MRIs and CT scans and those doctors have been pushed aside and they're gone. And you know, we get an mri, we get a CT scan, we trust pretty much that the images were put together properly and then hopefully someone read it right. You know, but that's the same thing. I think that's going to happen with generative AI. We're going to see where those edges are, we're going to find the areas where we don't trust and that's not what we're going to use it for. That's not going to be the value. There's going to be this area of trust and that's what we're going to use it, where we're going to use it. I think that those who use AI are going to be pushing out the people who don't use AI and that's how we're going to kind of move forward. And I mean, generative AI in this sense. Yeah, I have a slightly different concept here. All this time, as someone also said, whenever there's a technology, we just go blindly and start experimenting with it. Right. But in this particular one, if we drive through our use case, how we can use this AI to the best of our agency's use case, that will be better. So that's how USCIS we want to take. I've also had some discussions with Sean where we want to see if we can leverage what all are the tedious tasks today, especially in the uscis, for the adjudicators, the interview notes, they have to go through everything to make a decision, if that can be summarized. So that's one use case. So we can divide by use case and then you arrange the technology tools around it. So that would be even better for us to have a target. Because if we go through a technology, we are getting into what next, what next and we will still be in the catch up game. So. And just. Yeah, if I can. Just to add to that, as government agencies, we produce gobs of process and procedures, gobs of documentation. Those are prime for this technology and they're relatively low risk. Right. You have people who understand they're great for finding conflicts. Right. We've found security conflicts in our own documents using this technology. So great place to start, Great place to start to feel around for those edges. And prompt engine use cases are critical, for sure. Yeah, I love that. Can you do that so that all vendors can come through the door every day and. No, no, no, Meeting schedulers. Okay, got it. So what would you say now that you've talked about your architectures and where you want to go? What are the gotchas, what are the blockers? What are the lessons learned? Anyone you want me to take it? Okay. If I want to say from the lessons learned for past six years, whatever the journey we went, I think it's very overwhelming because every six months there's something changing. Even with databricks, we see this is introduced, that's introduced instead, as I said earlier, have your use case or what the program wants and then build around it that's more data driven or your program driven. Right. So that's how I envision which is more easier for us to transition. And if we go by that and then pick the technologies, it's easy. Now databricks is one thing, which is, which can integrate with any technology around, which is a good part. But at the same time, as a leader, if we want to enforce that first, we should at least educate the team. This is our final vision and this is what it has to be done. And then the technology should fall into its place. That's a lesson learned. And also take small steps because as soon as we start it, we think, oh, day one, we are here, in day 30, we will be everyone modernized. That's not going to happen. It's an evolution and it slowly evolves, but slowly at a point without even us recognizing we are already modernized. So that's how we experienced and now we don't want to go back to any of our legacy systems for sure. We kind of retired them. So we love that. Yeah, I think that. And when you say blockers, you're kind of asking, I'm assuming, kind of, what are some of the blockers to deploying new technology, to moving along this journey, this data and artificial intelligence journey. And I think that at least in my organization and maybe many others, and maybe this is a cop out answer, but it's kind of a balance of those four things that I talk about. People, process, technology, culture, each of those are all blockers. And as we are raising what I call data and AI literacy within the organization and we're involved in a technology revolution, we're involved in the largest acquisition, recapitalization of our organization and our history. As each of these buckets are filled, if you will, you got to kind of look at the other ones. The other bucket. Did I, what have I done for culture today? Are my processes now delayed? And I think that as you start to kind of put some focus on one of them and you aren't paying attention to the other, it becomes the blocker. So that's what I'm seeing as a cdao. Those are the buckets that I'm looking at every single day and trying to help the organization move along this journey. It's not just one. Yeah. And for me, it comes back to that. Simple, smart, reliable. Don't get over complicated with your architecture. Don't try to solve all your problems in your organization. Get started. Be simple at the start. Be smart about it. Pick the right partners to be with you. Make sure it's reliable. Make sure people can use it, and we'll adopt it. And understand that this is a military phrase, so you'll appreciate this one, Kev. No plan survives first contact. Right? So when you run into something, it's going to derail your plan and that's okay. Right. Get it back on track, make adjustments, move the plan forward. I think that's, for me, it's overcoming those blockers. The federal government has plenty of them. Right? We could sit up here for 30 minutes probably and talk about just contracts alone, about how they get in the way, but they should not stop you from doing this. They should not stop you from iterating, starting small and building to a large data platform. Awesome. Well, thank you so much for taking the time to participate. And we all know you're very busy and coming out to California and making the effort. We really, really appreciate it. So thank you. Thank you, thank you, thank. You.
First, thank you to all of our presenters listening to these conversations. I'm blown away by the initiative you're driving. While we're excited to provide the platform, it's the vision, the passion, the tenacity that has really brought AI and insight into your missions. And so I want to thank you for that. I want to thank all of you for that. Thank you. Now, this evening, the awards banquet, starting at 6 in theater one, will have the Award for Data Transformation Award, of which IFC World bank is a finalist, will have, Sorry, multiple pieces of paper. In this era of technology, we'll have Department of Veterans affairs finalists for the Data Disruptor Award, and we'll have the Australian Red Cross finalist for the Data Good award. So that's 6pm in theater one, but right now I have the pleasure of recommending the PUB SEC Industry Transformation Award. This is a first. This is a first I'm excited to announce. Veterans affairs is a recipient. As the FSC team makes their way up, I'd like to congratulate Veterans affairs the winner of the inaugural award with three brothers in the Marines. The mission that you guys are advancing is something of great passion for me personally, but also for us as a company. VA's goals represent really the best of the data forward agency. Come on up. There's room. We got four chairs here and come on. I think there's some more people from fsc. Yeah, get up here. They focused on leveraging the lakehouse for clinical, financial and supply chain use cases. We're excited to have many of the VA staff here from fsc, oit, vha, naii, EHRM and others. One of the best aspects of Data and AI Summit is a chance to gather notes to share best practices with your friends. Feel, feel free to grab them or any of the presenters here because really, that's the most important way we're going to advance the mission. Please take pictures of my grandmother in Puerto Rico. Yes, yes. For Vova or Nana in Puerto Rico behind the podium. Yes, yes, yes. That looks very legitimate. Oh, there you go. All right. Will you join me in congratulating VAFSC and all the great recipients? Congratulations. Thank you for your innovation. Thank you for your contribution to the mission. We have a lot to celebrate today, so we'll get a few pictures, but there's happy hour right out there. Lisa. Yes, yes, yes. This is again, a great opportunity for them to share ideas. I'm going to get off the stage so we get pictures of this great team. Thank you so much for your time attention today, we're excited to host you. I want more insights next year on stage.














