The public sector is undergoing a profound transformation, driven by an urgent need to leverage data and artificial intelligence for better citizen services, enhanced national security, and operational efficiency. Leaders from agencies like the FBI, Department of Defense (DoD), Department of Veterans Affairs (VA), and Hennepin County shared their journeys at the Data + AI Summit 2021, highlighting both the immense potential and the significant challenges in this digital evolution.
“Data is the new ammunition in our arsenal that improves decision making and connects the DoD's boardroom and battlespace.”
- Howard Levenson, GM, Area Vice President, Databricks Federal, LLC
Discover how top government agencies are leveraging data and AI to revolutionize public services, enhance national security, and drive efficiency. Learn about the critical challenges and innovative solutions propelling the public sector into a data-driven future.
hello everyone thank you for joining our government industry forum at the data nai summit i'm michael ortega from the industry marketing team here at databricks and i'll be your host today we have an amazing group of government leaders from some of the most respected agencies joining us here today to share their perspectives on how data ai and the lake house architecture are enabling them to improve citizen services better protect against threats and deliver on their mission objectives without further ado i'd like to hand it off to our first speaker today howard levinson welcome my name is howard levinson i'm the general manager for databricks federal and i'd like to welcome you all to the public sector industry forum we're gonna have a great set of chats today about uh accelerating ai initiatives in the government space and the goal is achieving mission outcomes so let's get right into it so we've got a great panel today there's actually two different there's a keynote and an expert panel and i'm going to hand it off after this brief introduction to john larson from booz allen he's the senior vice president at booz allen in charge of analytics and on his keynote he'll introduce depali pawali who's a data scientist with the federal bureau of investigation as well as honorable greg little who is the lead for the dod havana project following that i'm going to host a panel with dave fuller from veterans affairs paul donahoe from center for medicaid and medicare services and andrew lucius from the hennepin county minnesota so let's get right to it and uh let me give you a little bit of background about what we're going to talk about today so we know that data and ai is a top
white house priority we've seen a whole bunch of legislation to kind of drive these points forward we've got the it modernization act and the goal there is to move from legacy on-premise data centers into the cloud we've got the federal data strategy the responsibility there is to federate all the data and catalog all the data and governance all the data govern and govern all the data so that we can use data as a strategic asset and finally there's the ai executive order and the goal there is to drive data and ai to making evidence-based decisions it's not only that there's a lot of legislation there's a lot of money behind this as well government last year spent a billion dollars and will spend more than that in 2021 to accelerate data and ai research when you talk to federal cios out there like gartner did you'll find that the top priorities are none other than ai and ml machine learning data analytics and cloud and the justification for all this investment and the costs of all of these priorities is tremendous outcomes for our government and for our citizens in fact deloitte estimates this to be 41 billion dollars in savings every year from data driven automation but there are challenges in moving
forward with this process first off uh we've got a skills challenge you know it's hard to these are new technologies and new capabilities and frankly the public sector is behind the industry average in terms of the number of data scientists and data analysts that they've got so we've got to close the gap there we've also got some gaps in the it modernization act you know a couple of years ago opm did a study and they found there were twelve thousand data centers across the federal government so it's going to take us some time to migrate that data into the cloud and it's hard to do that moving if the data is all isolated in on-premise data centers you can imagine how difficult it is to build a federal data strategy with all of these islands of data and we all know that you can't use ai and ml until you've federated all of that data together and you've normalized all the data before you can actually start applying advanced analytics so how do we resolve these challenges
well the way the approach that databricks has is this lake house pattern in this lake house platforms and so let me describe the problem a little bit more and then we'll talk about how lake house addresses it but the challenge again is you've got all of these data sources that you see on the left many of these are operational data stores in relational databases that the federal government's had for decades and the idea is we want to take all of this data and we want to federate it all together so that it's ready for analytics so the goal is you take all of this data and you migrate it into the cloud and the cloud offers this really low cost durable storage that's highly reliable and you can put all of that data there at a very low cost once i've got the data there and i've normalized it then i'd like to start doing analytics on it and hopefully from those analytics we'll be able to deliver better results that may come in the form of better national security in terms of reduced financial crimes better social services and many other capabilities the government can benefit from better use of data so let me show you how we do this with the lake house platform so first of all databrick started as a cloud platform so we've always been native to the cloud and the idea here is you take all of these operational data stores and through tools like our delta and our delta live tables you can take the data from those operational data stores and as the data is changing you can migrate that data into the cloud through a process called change data capture and this is enabled in an automated process through delta live tables and the underlying delta file format once you've got that data in the cloud now i've got it in a low-cost highly durable cloud object storage and i can normalize it i can get the data to look the same and act the same and now i can start to run analytics across all of that data and it's the databricks delta file format and delta engine that enable you to do this at scale but we also offer a managed catalog that allows you to govern this data and track the data so that you know what data that you've got and you can make it available to those people and those personas within your organization that are entitled to it from there you then build and train your machine learning models or you can bring your sql analyst directly on top of this data without having to first migrate that data into a data warehouse so you've got this analytics capability that allows you to do everything from sql analytics to machine learning and artificial intelligence all built into the comprehensive system ultimately providing you a bi machine learning without any limits now let's go through that in another step so the lake house architecture is really the best qualities of a data warehouse which is the ability to deal with sql analytics and with structured data along with the best qualities of a data lake which allow you to deal with all of these data types and allows you to use advanced analytics on the data and not just sql in the databricks lake house paradigm we bring both of these together and of course we run on all three cloud service providers azure aws and google all of the data then gets stored in the open data storage which is our delta lake capability and from there you run our manage catalog to deliver the management the governance and the catalog so that you know what data you've got there on top of all of this we deliver a platform that's ideal for all the different personas that deal with data in your organization data engineers who are responsible for ingesting the data transforming the data and making it ready for analytics bi and sql analysts who want to be able to build uh standard reports on that data data scientists and machine learning engineers who want to be able to do predictions and determine anomalies from the data and even real-time data applications that can take iot data in and actually uh perform analytics on that data as it's flowing in now all of this runs in the databricks platform making it a simple platform for all of your workloads it's an open platform so there's no vendor lock-in almost all of the capabilities that data ha databricks offers here including spark including delta and ml flow have been open sourced and it's a collaborative environment with the ability to share information with role-based access controls so that only the data that you want shared gets shared now the other challenge that many of our federal clients have is that they want to run this in federally compliant environments and you shouldn't worry about that databricks offers uh this capability for dod clients for intelligence community clients and of course for civilian agency clients as well as state local and education uh organizations so we meet just about every criteria you could imagine including um you know fedramp for those civilian clients fedramp high for even more civilian clients we are we meet the uh dod's uh il-5 and il-6 requirements we even meet the intelligence communities requirements based on icd-503 now you don't have to believe me on these things we're happy to prove it pick a use case bring some data bring some of your subject matter expertise we'll enable databricks in the environment and we'll bring some of our subject matter expertise and very quickly we can build some analysis and show you insights in your data that you've not seen before and more more importantly we can show you the business value that you can get out of that data so i hope you'll join me in the conversation but at this point i'd like to turn it over to john larson from booz allen as i mentioned john's a colleague a good friend and the senior vice president for analytics at booz allen and he's got an interesting approach on building a standardized platform for doing big data analytics here take it away john
thank you howard for that great introduction i'm excited to have this opportunity to speak about some of the innovative things that we're doing in partnership with databricks on behalf of our clients most critical mission challenges by way of brief introduction my name is john larson i'm a senior vice president with bezalen leading our ai and ml solutions for our civilian and commercial clients today along with depauli pali a data science lead and program manager with enterprise analytics at fbi and greg little deputy comptroller for enterprise data and business performance at the dod we will discuss the challenges and the solutions to achieving enterprise ai as i mentioned in my intro and if we could move to the next slide i'm with booz allen a leading provider of consulting services and technology services for the u.s federal government and fortune 500 companies we were founded in 1914 and today we're the largest provider of public sector ai services with 30 percent of the market share we have over 27 000 employees and a critical mass of 500 plus seasoned ai and ml practitioners and we're recognized on an array of awards but i'd like to specifically highlight too the 2019 inform prize for innovation and analytics and in 2020 21 cms ai health outcomes challenge where we were part of the winning team so we do know a thing about ai and how to do it successfully with that let's go to the next slide and talk a little bit about why is ai success so hard we've heard all the stories we've seen the headlines achieving ai in a lab where it's sterile where the data is clean where the processes are well controlled is actually relatively easy it's when you scale this and the applications and bring them into the production environments with massive and messy data expensive compute resources and complex business processes that things get difficult in fact the vast majority of ai and ml projects will fail because of these challenges and they will never move to production with more than 300 concurrent analytics projects we've learned that there's really four key elements to doing ai successfully in an enterprise the first is cultural readiness the second is data readiness the third is technology readiness and the fourth is process readiness and today along with the deep dives from depaulie and greg we're going to focus on the insights around achieving data and technology readiness and how we accomplish those through the databricks platforms so if we go to the next slide we'll talk a little bit about the specifics of our extensive experience and what it's taught us as we've worked with our clients to address enterprise ai and what we've seen is the standardization of operations around these two pillars of data and technology is critical ai technology is simply too large complex and rapidly evolving to make it very cost effective or difficult to stay up with in terms of all the possible permutations that you have to govern so it's very critical that you think about those different solutions and how to integrate them together with the appropriate architecture it's also difficult to obtain and scale talent the skills required to bring these different tools and techniques together is very difficult to maintain and acquire and that means that you've got to have robust ai solutions that have a consistent and standardized architecture to give you the best ability to scale your solutions and finally while i think everyone looks at all of their problems and believe that they are truly unique the reality is is that the foundational concepts remain the same throughout most ai applications and so you have to find that consistent foundation and extend it so that you have the ability to create that scale more rapidly and we found institutions really do increase their probability for ai success when they make it easy to apply and reuse these standards by defining common terminology to make the architecture easy to consume regardless of the experience and if you focus on democratizing ai and creating an inclusive community with seamless collaboration tools to unlock the potential and the creativity of the workforce you also increase your success rate and finally if you drive constant learning by leveraging evolutionary and open architecture frameworks and embrace change you increase your probability of success so that's what's really driving the keys to success here and if we move to the next slide we'll talk a little bit about the crosswalk between what we developed in our reference architecture and what we saw in the databricks platform and from those lessons learned we developed an overall reference architecture that looked at how do you leverage established providers and services rather than reinventing the wheel this is an expensive space to play in and it's generally better to integrate together best solutions from others than trying to build your own we also found that everything is evolving rapidly if you blink you're going to miss something and so you got to stay open architecture you got to stick with the evolutionary open architecture approaches approaches that embrace change rather than resist it and evolve with technology and client needs we also found that you've got to create a suite of collaboration and productivity tools driving collaboration across data scientists data engineers ai and ml engineers requires notebooks and github repositories integrated into the platforms you're using to ensure collaboration and achieve the scale that you need and finally we found that if you drive smart automation constantly train and use partnerships to drive that scale that's a much better architecture to hang from than if you don't adopt those types of techniques and so those principles right as we looked at those from those principles we created this reference implementation within databricks databricks because of how closely their platform aligned with our reference architecture and what we are trying to do in that vision and what we really enjoyed about the lake about data bricks was the lake house architecture which enables reliable data centralized for analytics and ai workloads there's also an incredible interoperability to support the scale and operationalize aa across a variety of use cases this is how databricks really enables integration with reusable ml pipelines to minimize variation in complexity and costs how they give end users the control and choice over applications of ux and ui layers so that you don't have to do rework because it's already built in and how they enable collaboration power with notebooks ci cd and github we also appreciate the automation i think it can't be understated how important it is to have auto scaling auto termination and performance optimization built in these are significant cost savings and we found tremendous value in the life cycle machine learning governance and tracking and how it offers configuration management of data features and models and it decouples the machine learning software baselines to deploy models faster more frequently and independent from software to enhance model performance and adoption and finally i think the other thing i will just emphasize is back to this notion of speed the wonderful thing about the integration of these solutions is that it achieves speeds two to ten times faster in terms of its optimization around spark to significantly reduce csb cloud saving costs so that you can enjoy higher savings and so that's why we looked at the databricks platform and said our reference architecture needed to be developed and deployed within this framework and we developed this reference implementation with that in mind so with that sort of overview what i'd like to do is i'd like to turn it over and say let's have a deep dive with the poly around some of the things that she and her team have been working on and then from there we'll have a conversation with greg about some of the applications that they've been deploying so with that let me turn it over to depauli hey john thank you so much for the introduction and a good overview
i think i'm going to touch on some of the points that you have already kind of covered and how we are leveraging in our data pipelines so i'll talk about our project prometheus it's our content warehouse where we have leveraged data bricks in many of its components and data in just pipelines we started with a data ingest i'll talk about that next about our enrichments later about analytics and the sandbox environment how we leverage the collaboration with our mission users and the enterprise we'll be they'll be seeing some of the use cases about that so as you all know our data ingest and data is very complex our data pipelines are very have to cover a lot of a diverse nature of the data along with apply the security policies and compliance we have a user-based user attribute based security along with law enforcement policy-based security so this makes a security model very complex to apply than the standard data processing pipelines so most of our solutions was based on nosql database most specifically which allows us to do cell level security since we onboarded databricks in 2019 q4 we have simplified our data pipelines our ingest pipelines where we have the reliance on nosql we have minimized our reliance on our nosql solutions and broke it down into what has enforced is a data lake house now but before we started with a delta lake kind of implementation this has significantly improved the performance for us it has many of the siloed components were but no more needed we have also delivered the delta lake formatted output which is now applied to many of our downstream stusk as you all know the delta lake also provides us an asset compliance and the revisions and the history which helps us to leverage some of the a complex in in just pipeline related stuff so for example if data needs to be re-ingested for due to the policy or due to the any other changes in the data we could get the latest and greatest revision in the analytics pipeline so this kind of easy use of the data was never was not possible before another use case is how do we purge data when there's a there's some kind of policy changes that perching in the one side of the pipeline should not affect the other side of the pipeline so having using the revisions making sure that the data is compliant from end to end was of a great use that we have seen today we are using data lake uh data ingested data from the raw data that comes up from our providers and we are now able to leverage it in our data lake pipelines uh that's helping us to kind of build that end-to-end discovery and end-to-end um understanding of how this data can be leveraged using a delta lake house now next thing what i would say is that because of the simplification of this end-to-end pipelines has helped us to focus more on the data discovery building an apis to discover this data so has freed up a lot of our time with our data engineers from the analytic side of the house we were the early adopters of spark so moving to databricks was no brainer we were able to convert all of our analytics pipeline into databricks in almost a week's time we were able to find the performance improvements the jobs that were taking many hours were now finishing in in minutes that definitely got attention of many people we were able to kind of have build a data analytics pipeline given in the same format which is a delta link format from our ingest we don't have our simplification of a data ingest or data analytics ingest pipeline was very helpful here we could focus on enrichments building data sensing techniques feature engineering rather than making sure that the data is compliant making sure we are uh we are not making any policies and those kind of stuff was happen because of the data breaks uh in just pipeline other part i would say is that spark and processing using spark for building up our aiml use cases especially the machine learning use cases building our own algorithms was very helpful applying mathematics complex mathematics is required for some of our use cases having the processing done in memory in a distributed environment we were able to get the slas that we are required for our downstreams task the other reason to use case that i can say is that since 2019 because of the simplification of the pipeline we were able to release our a big analytic that leverage leverages data mining techniques record deduplications graph analytics to the enterprise and that has been a very successfully uh i would say visualized understood by our uh our end users our enterprise users and we were able to create the apis that can be integrated in our um sandbox solutions which i'll talk a little bit shortly another part that we have also done the
most recent work is building ai proof of concepts as john was just talking about sort of like building that ai frameworks or ai philosophy it requires some sort of an adaptation some requires a baselining some of the work that is required to prove out whether this ai concept is going to work for you or not what we have done here is our ai piece was mostly doing an extractive summarizations that are that draw insights from our analytics so here we were able to use the databricks end-to-end what i mean by that we were able to pull the data from our uh data ops pipeline which was not complete but which was done just about to use in foreign analytics we pulled the data from there and reached it make build useful analytics applying complex mathematics on it and then delivered the analytics which was visualized using open source libraries and built an extractive summarization component of it so all this was done using databricks visualizations were also well i would say well received by our ux team we have to demo this to our executives and one of the thing when this project was kicked off i told my team that we are not going to use powerpoints we are going to use this delivery end-to-end using databricks and let's see how what are the other things that we fall short luckily we didn't have to use any any uh powerpoints so that was a one big huge benefit but the idea behind behind not using the powerpoint was that to explain the dynamics of the mathematical models that are being constructed can we effectively do that explanation can we show the different steps that goes in to build that models can we add in the explainability over there so some of those were the considerations for us when we were building this proof of concept and i don't think anybody was saying anything was anybody that was missing any of the powerpoints so was very successful in my opinion also the same analytics visualization is not been kind of given to our ux people who can actually use it to build it in our enterprise uh uis next i'm going to talk to you about the our sandbox now that we have this data at an enterprise scale and an enterprise level which is kind of stowed in our data lakes and content warehouse we wanted the users to start using this not just the enterprise users who are technically savvy but the users from the field who have a multitude of experience of knowledge of how to use the data they might have done some python development they are sql experts but they are also citizen data scientists who we want this data to be exposed to so we were looking for an environment which was easy to access at the enterprise scale they should able to access the data that they have access to so there has to be a security-based model they should be able to search only what they can see they can check out this data and also bring in their own data to kind of pull this information together and build the analytics that is of interest to them at the same time we wanted to expose all these analytical apis that we were developing the analytical the enrichments that are being available to them through this rest apis that can be incorporated in their work so we were looking for a platform and databricks fit that bill at the same time we wanted to reduce the cost we wanted to make sure the compute resources are not overly used there's a proper governance on it at the same time we wanted to make sure that our clusters are shut down at the specified time and and people are no wise users should be able to feel comfortable using an environment and that's why databricks was one of the environment that we chose uh for the sandbox now what another thing we were also looking for was the self-paced training databricks provided us with some self-based training they should also provided us the rsas where they can work with our sandbox users can work with these rsas feel comfortable to ask them a question and kind of really literally leverage some of the enhanced capabilities of databricks in their use cases another thing i would say is that what is very working in this environment is a collaborative nature people should be able to share the notebooks should be able to do the demos should be able to kind of download the data of the insights that they're they're building in this notebooks and can and kind of use it for their um their reporting or their work that they do for the mission what we have seen is that the users have not only not only leveraged all these different features they came up with the use cases different many different use cases that as an enterprise analytics we have not even thought about and really solve them using some of the analytics api that they have access to some enrichments that haven't built and really literally taken it to the next level for them to solve their mission use cases later we also have noticed that the users were also not just this we started with a very i would say about dozens of users later we got some more requests more use cases to solve uh finally we would have to say like because of the of the nature of this uh this delivery we have to kind of kind of open it in a later kind of have a time box this uh user onboarding and what we have noticed is that people's enthusiasm to use such sandbox environment was was amazing i would like to talk about how we can leverage this feature or this kind of collaboration with the enterprise users who has a multitude of different skill set with within analytics with within analytics or especially aiml workloads many of the insights that are drawn by this these users are can be combined to apply to an enterprise services where users should feel very comfortable to to to input or rather to provide their inputs and have their say so something in in uh when we build over any of the analytics using an agile processing a story point that is or the story which which says that hey go and build this analytics and then has a completion criteria we have to go beyond that we have to go think about at least four to five different steps ahead and these users are going to help us to do that rather than having a third step for us a third model that will help them use this use case we wanted to early on capture that requirements so i see this as a collaboration from all different types of users all different skill sets people who are are citizen data scientists but people who are also skilled uh data scientists many of our aimo experts can also work together on this and this end-to-end pipeline where you are focusing on analytics and rather than um rather than thinking about how to manage the data is huge it's a huge help and huge um i would say a step ahead for us this is how we leverage data bricks uh back to you john yeah thank thank you to paula and what a great uh set of use cases and thank you for sharing that it's really exciting to hear you talk about sort of the promise i feel like when you and i were talking earlier at time immemorial it's always been this data challenge you know you spend all your time on data and hearing you talk about like we're realizing the ability to spend more time modeling feature engineering tuning like that's exciting that you aren't just trying to figure out and spend all your time in the in the data right and then i love the other example around this citizen data scientists and these sandboxes and distributing this sort of capability and leveraging across all these let's use the term sort of nodes your your sort of distribution center if you will through these collaboration tools and driving greater insights it's really exciting to hear that come to fruition and see the type of power tools like this can unlock for us so thank you again for sharing that with that let's pivot let's come over and talk with greg a little bit who's going to talk about the advanta platform and some of the things we're doing around the application of this technology stack in the dod environment so greg
yeah john thank you for the invite to discuss dod's data journey and depauli i thought your presentation was incredible and what it did is it really highlighted an organization that's really mature in their data space and can actually use ai and ml well i'll talk a little bit about ai ai and ml a lot of our journey is still around the data and getting the data right and i think we all know that if we don't get the data right using tools like ai and ml are almost impossible and in dod we really view data as another weapon in our arsenal to compete and provide strategic advantage and as my friend dave spurk who's an incredible colleague and the dod cdo likes to say data is the new ammunition in our arsenal that improves decision making and connects the dod's boardroom and battlespace and i know that we don't have that much time today so i want to quickly share our story around how the financial statement audit fundamentally changed how we think about data in the dod which you when i when you think about it'll actually make a lot of sense but compliance was sort of our entry point into the data space and the analytics space and part of that was the creation of advance which is a program that i own which is dod's enterprise unified data management and analytics environment which is made up of a collection of open source and cost technologies and then i want to highlight some of the struggles that john talked about today and share some of the successes that we've seen within dod especially around using data analytics to improve dod's mission performance so i currently work in dod's financial management community and this idea of using data for better decision making is not new if if you actually look back at the cfo act of 1990 the whole intent of that act is to have better data to run a better government and so just to give a sense of how long ago that was i looked up last night that in 1990 the number one song was ice ice baby by vanilla ice and the number one movie was home alone so if any of you want to feel really old look up a picture of what mccully culkin looks like today and you'll feel extremely old so knowing how long ago we've thought about using data analytics to improve decision making the question always comes up is what's different this time because i think there's a heavy amount of skepticism within the community around how we're going to actually realize this vision of using data in analytics and ai and ml to improve performance outcomes and i think there are really four major differences of why i have hope for dod the first is leadership and we're really lucky to have we had a good leadership in deputy secretary norquist and really incredible leadership and deputy secretary hicks in this data space you just signed out a memo called the decision advantage data advantage memo and in essence that memo really elevates the role of data in our governance and how we're going to run the department of defense but not only is she issuing policy secretary hicks is actually walking the walk and we actually held a performance review using advan our enterprise data management analytics environment looking at how the health of our financials the health of acquisition supply chain operational readiness and we did this with live data and no powerpoint which is a small miracle within dod for sure the second area that i think is really different this time is around technology now i always have a cautionary tale here because i think oftentimes we confuse our data strategy with a technology strategy and don't get me wrong technology is an important part of realizing the power of data to improve outcomes but technology is just a tool and there's really an ecosystem that goes around technology to be able to make sure we're using in the appropriate way to actually get to the outcomes that we're looking for now with that said there are three major technology shifts that have really accelerated our data journey the first is the cloud the scale flexibility security and the price of the cloud has brought speed and scale that was not possible even five years ago so for example at the rate and velocity and volume of growth that we were seeing within evanna in order to plan compute and storage out on premise would take me months and now i can scale up and down within minutes the second is that the tools for data engineering data science and data visualization have become easier and easier to use and this has really allowed our business analysts with little to no coding skills get insights from data with ease and so technology has really allowed me to tap into an additional pool of resources that probably weren't possible five years ago and that's something that we could really use help with our vendors is to be able to make these tools easier and easier to use the third is schema-less technologies and depaulie mentioned the nosql databases before when i when i first started out in this space uh we had to spend a lot of time on etl we had to spend a lot of time on getting our source systems to provide standard schemas and with new technologies today we've been able to sort of take in any format any schema and ingest that and that's really saved in my opinion years in us being able to get access to data to not only be descriptive but but to also be be predictive and how we're thinking of data and so in my mind the bottom line is that technology improves speed and speed really matters when you're supporting a mission like dods now the third thing that i think is is different this time for us in dod is the audience i sort of mentioned earlier how the audit has actually become a surprising story and accelerator for us in our in our data and analytics journey and the reason why is that the financial statement audit has been one of the best data management forcing functions in dod for having an enterprise cross-functional and organizational data management analytics environment my first requirement when i started advana which has become that enterprise analytics environment was to deploy what auditors call universal transactions and really a universe of transactions is simply to show the auditors we can reconcile all our business transactions like contracts invoices inventory property personnel to our financial statements now while our financial statements are not particularly useful for dod that transactional level business data is and also the principles of audit are the same for good data management we want accurate authoritative complete and timely data in addition the audit broke through some of the cultural barriers of data fiefdoms when your large organizational bureaucracy like dod you need to really look for institutional levers that can break through cultural and organizational resistance and sometimes compliance is the best way to start an initiative like this so for example the audit occurs every year it'll never go away and everyone has to participate and there's really a non-controversial ad proposition people can either give me all of their data and i can work with the auditors or they can have the 400 show up at their doorstep and for most of the dod community they'd rather have me deal with the auditors and give me the data which opens up a lot of interesting analytics for us to be able to to get to those outcomes that we're looking for in the fourth area that i think is really different this time is that oftentimes people like me and dod will stand up from the ivory tower and talk about theory around data and how we're going to improve performance but what's cool about about over the last couple years is that we're actually doing things in this space as i mentioned advana is that enterprise data management analytics platform we have about 18 000 users we ingest data from over 250 data sources we have about 25 billion transactions and we grow almost 500 million transactions a month by having and starting out with the audit and having all this cross-functional and cross-organizational data cross-organizational data we were actually able to respond to the covet 19 crisis by building a coven 19 common operating picture for dod in less than three weeks and this common operating picture was looking at bed utilization we were running models to be able to predict risk where we had to move certain pppe like ventilators and we've been using it to look at our vaccine distribution metrics and determining where we should move vaccines to our higher risk populations and so i'm really proud of what we've been able to do there with with data using some pretty cool analytics techniques we also are working with a lot of our senior advisors in the department on some really important strategic initiatives like diversity inclusion and climate change we've been able to bring in all the weather data so hurricanes fire flooding map that to our installations to look where we might have operational vulnerabilities due to climate change we've actually used analytics and decision making meetings as i mentioned earlier the deputy secretary secretary hicks is actually using advana in her decision meetings in our business health metric reviews using live data and no powerpoint to help run the department of defense which is a really a 3 trillion in asset enterprise and we started to move not just looking at descriptive statistics but understanding what happens when we can start using ai and ml so we actually have our acquisition sustainment partners looking at our major weapon system programs like the apache and the blackhawk and actually using ml and ai to predict the health of those programs 60 90 120 days out which is which is pretty amazing to think about so we can forward plan and find areas of risk before they even happen the other thing uh that people might not know about where we've started using some different analytics and aiml techniques is that the federal government actually has expiration dates on its funding so some money you can only use for one year somebody you can only use for two years and so before that money goes back to the treasury we want to make sure that we use it so we're maximizing our better buying power we've been actually able to use some interesting algorithms looking at certain dormancy of our transactions being able to rank by dollar amount and other factors and have actually been able to repurpose almost four billion dollars to high priority items the last one i want to just share real quick in terms of i think a real success story and the financial area is really robust in terms of ai and ml and pretty mature is that we've been actually able to suck in all of our payment data for the department all the all the payments before they actually occur and run some interesting algorithms that have actually been able to prevent almost one billion dollars in improper payments and so what's been really cool over the last two years is that we're making this a real reality now with some of these successes i don't want to give the impression that we're not just starting out i mean we have a long way to go to really get to the nirvana state of using data in every meeting at the strategic level in the operational level to help improve performance and run the department
but we know that data is the key to success in this digital age and like many of you we in the dod are still struggling to redefine our strategies around data and for me and john talked about this the battle or at least the discussion has centered around two approaches one is tearing down data silos and the second is creating a data-driven culture i think we've really struggled on both of these fronts strategically the data silo accessibility struggle has somewhat missed the point in my opinion yes that data that is locked in one application or database inaccessible to other users is a problem but the key to innovation for us is really when we can bring data together from multiple different functional areas like acquisition and tied to financial management and tied to it or tied to readiness and i think the one of the things that we have done wrong in dod is that in order to do that and unlock this data we try to make everyone be standard and i just don't think that's possible or needed in the technology space i actually think it'll hurt dod because the database and structures that those source systems reside in today are done in a very purposeful way for that particular use case and so we really need to be leveraging technology and a new way of thinking to be able to tap into those source systems to be able to take the schema as it is and then to be able to bring into a platform like advana and use technology that allows us to connect those functional dots and not be encumbered by that primary data source or the primary organization of responsibility so it's really freeing up data and having an open data architecture is something that we're really working on in dod probably the biggest challenge that we see that john also mentioned was culture in my mind a data-driven culture is one in which data is used to inform every decision in which every action is a strategic choice using the best available insight in the tools that provide insight lead to the tools such as orchestration automation that drives swift and effective action and you can't use data to drive every action until you give every decision maker hands-on access to data and the tools to act on it and so as i tell my team it's not really a chicken and egg question in my mind the business questions and the data must become before adoption and we've been working really closely with our functional areas like i.t financial management acquisition to get the right questions the right data and then use the right tools to answer them and so let me let me wrap up with this i believe this combination of an organization-wide platform like ivana in a culture that embraces it is how dod will compete in this ever-changing world environment we need to guide the evolution of technology and culture in parallel so that we can use data across our organization to discover new insights that guide dod strategy ignite innovation and really redefine how we achieve success and so with that john i really appreciate the opportunity and i'll turn it back over to you great yeah thank thank you greg really good insights and i think um you see this sort of the the nexus between these two use cases i heard a commonality around the speed to insight from data the complexity of overcoming those data challenges and how you can use this technology to accelerate that i also think that you know you said something that was very consistent with some of the the the insights the poly that you raised which is this idea of how do we abstract out and and step back from all the technology and allow sort of a broader consumption of citizen data scientists um you know business analysts to be able to use low code no code or other types of tools and techniques more easily without having to understand all the back end and create and drive insets and insights and i think the last thing i think you said that was really compelling was i think both of you are envisioning a world in which powerpoint is not the primary format of data consumption and insight which is i think something we can all align around is the idea of how do we become data driven and how do we look at more dynamic insights as we sit in meetings and drive insights and and actions from that data so that we're making uh data-driven decisions i think greg you know you talked about the the data acts of the past i'm thinking about the evidence-based policy act and and its importance and the role of making data help inform every one of our policy decisions so really great conversation thoroughly enjoyed it thank you both so much for the time and for your commitment and leadership in driving data and analytic insights for our government and for the success of our nation so thank you very much and thank you databricks on behalf of of booz allen for allowing us this opportunity to talk today thank you thank you john depauli and greg for sharing how you're driving innovation at the dod and fbi powered by data ai and databricks i'd like to now hand it back to howard and our esteemed group of panelists for today's government panel discussion take it away well hello my name's howard
levinson and i'm the general manager for databricks federal and i'm really excited to host this government panel today today on the panel we've got dave fuller chief of the data analytics division of the u.s department of veterans affairs we have andrew lucius who is the principal data scientist from hennepin county and mr paul donahoe who's a senior technical advisor in the enterprise database group at center for medicaid and medicare services so why don't we get started just by a brief introduction dave maybe you could start off give us a little bit more about yourself where you come from and uh what's uh what's going on at the va right now great so i have the distinct honor and privilege of serving our u.s military veterans and their families through supporting the department of uh veteran affairs program and operational decision making my chain of command includes the office of management and the financial services center so i oversee data analytics products and share the responsibility of architecting our devops pipeline and migrating to the cloud and dave how long you been doing that for uh three years three years very nice okay let's move on to andrew andrew give us a little bit of background and the role in your organization and what you're doing with data and ai sure so uh i've been with hennepin county for five years i'm a principal data scientist as howard said and my job is at least currently split between three areas so i work in part in the county attorney's office in part for public safety administration which sits sort of above the county attorney's office and all the aspects of public safety and then i'm also on the covet data team right now which is sort of a temporary emergency team to help with all of that data dave i'm sure you're getting three times the pay these days not quite but and over to you paul i know you've been doing this more than three years give us a little bit of background on the work that you're doing at center for medicaid and medicare services and who paul donahoe is more than three decades actually howard i mean i mean 35 years almost all in uh big large systems with data so i've been following the data uh you know evolution as an industry for uh quite some time so uh been at cms for 22 years uh most recently as with the marketplace helped uh work that system there with the affordable care act uh recently i'm working out of the office of information technology as a data strategist uh we're trying to figure out you know what's what comes next now that uh there's a new generation of tools and uh methods out there so we're trying to figure out what's next you know i love talking to you paul because you're more than just a data scientist data analyst you're like a data historian and you can tell us how cms has evolved over the decades and why we are where we are today uh as a result of what's happened in the past so i always find that fascinating hey you gotta know where you're coming from to know where you're going it's so true man maybe you can expand upon that since you got the microphone tell us about kind of like the analytics journey at center for medicaid and medicare services kind of where it started from and then maybe if you could also share how it's uh better or different from some of the other federal agencies that you've interfaced with well first of all i love the the mission of cms it's just uh it's a great place to work it brings a lot of talent to you know so anyhow it's a really great organization so if you're out there you're looking for a job cms is a great place uh it's it handles a tremendous amount of data uh well we impact 17th in the united states economy and so that's fixed so it's not just health security it's uh you know uh financial security for the nation as well so but you've got i mean just look at the programs uh children's health uh medicare medicaid the marketplace each of those systems have tons of data uh you know i saw it through where we had mainframes and tapes and just mountains and mountains of data uh and now we have many more sources but you know we're moving to the cloud and i guess the big the big thing for me is you know i was there when we went from uh hierarchical to relational from um mainframe to the mid tier but the biggest change i've seen uh in my career has been the uh separation of compute storage now i know that gets a little bit squirrely to think about but in the old days you know you had to have the compute the memory and the storage all sort of tuned up and i couldn't let you if i'm on a project i couldn't let you if you were on another project on my system to share my data because you might have a runaway query and that would mess up my ability to pull off my service level as a project manager well when you get to the cloud and then you know that's how we've uh learned of each other howard is the uh spark is the first you know many many tools come out today uh that do similar things but but spark is uh was that first i don't know miracle came out of berkeley that separates compute and storage where i could have the data sitting over here and many different teams tapping into it neither uh knowing that the others there and they're not degrading one another and uh boy that's a big thought and it's going to take us a generation to figure out how that all really comes to play but i've got some ideas on that but we'll let some of that go for a little bit later in the uh discussion here howard but i think that's the big ticket is the separation of compute storage what does that open up uh in the cloud in particular with the external compute it's just a brand new ballgame and it's very exciting it's a very very cool thing but anyway and i gotta hand it to you paul because you recognize that differentiation before most people in the industry did and kind of seized upon it at center for medicaid and medicare services so we're going to come back to that because i think that's an important theme but uh just moving along uh andrew lucius maybe you could tell us about the analytics journey at hennepin county and i'm particularly interested maybe in if you know since it's uh so relevant uh what's happened with the covid use case that you referenced earlier sure yeah um they're sort of parallel stories i suppose so with covet data you know it was it just hit everybody you know at like a train all of a sudden we had to get a bunch of stuff together very quickly um for us at the county level you know we needed to run isolation and quarantine sites for the homeless population get those up and running very quickly we need to get surveillance tracking going very quickly we have a public health department within our county and so the chief data officer just pulled in a couple people because he was getting you know that pressure from above to get reports together and we just we like r we're our people um and so we just set up a ton of our scripts basically that that ran locally on our machines off one drive pulling you know because everybody's scrambling so a lot there's a lot of excel sheets flying around and we were just pulling things together um putting out power bi reports trying to track everything that was going on trying to clean everything in our and it had us going you know seven days right so every day we'd have to do something manually on our machines and things would break constantly because you have the spreadsheets and so databricks came in for us as an automation tool essentially so we really didn't need the big data aspect of it i know that's usually what you hear about with spark but the nice thing about data bricks is you also have that single node cluster option and you also have you know python and r preloaded on there and most of the packages you'd need to do basic data cleaning stuff like that and so it was really nice for us we could just move our spreadsheets to the lake with a logic app um and then we could automate what we were doing with the the cova data cleaning and data bricks on single node clusters and gave us our weekends back essentially so um yeah and as an organization um we're still thinking through what we're going to do with sort of enterprise data warehousing strategies but within public safety the project i'm on sort of got that ball rolling earlier we're trying to get all the data from from all the public safety agencies that impact hennepin county that interface with it so inside and outside sort of into one place you know and a lake is obviously ideal for that that's a really key point that i'd like to come back to as well you know the i just was reading the dod has their own data strategy and uh you know really paramount to that is uh data is a resource and it's got to be federated in in order for them to really make valid get value out of it so uh we'll we'll come back to that and talk a little bit more about uh you know all these different islands of data and how you pull them together
hey dave tell us about the analytics journey at the va you guys certainly have tons of data and uh i know that at the beginning when we got started with you guys you were largely on premise tell us a little bit about your your journey right right so that began about seven years ago we had just a few resources you know a small contract and big ideas on how to support the regional cfos across the va we picked up some of the customers that manage the healthcare supply chain along the way but since those early days so much has changed now we have grown our federal staff increased the size of the contract expanded our customer base and even continue to have big ideas on how we could support the va leadership we want to make sure that they have some kind of efficiency to accomplish their mission of serving our veterans right so our particular has the quantity of data grown considerably oh yeah yeah for sure um but before i get to that you know our organization i you know yes how we're different um we're a franchise fund and that means we're not funded like you know the typical government organization instead we operate more like a business we're funded by the agreements that we make with other organizations across the va or in fact with any other federal agency uh so we have to pay real close attention to our you know our revenue and our expenses but we're pretty decent at that because we're also providing financial services dashboards across the va and you talk about large data one of our use cases that we'll talk about is providing an enterprise solution that's aptly named the cfo dashboard the data source for this is the financial management system that contains the financial general ledger for the va and you know that's significant because the size of the va budget is only second to the department of defense right wow so yeah the cfo dashboard it basically provides uh summary and detailed drill down capabilities with various visuals to understand the status of obligations accounts receivables advances and we even added you know you were talking about covet 19 we added cobia 19 spending tab to accommodate that current focus that yeah that we have that's pretty cool that's uh really cool i know you guys have a massive amount of data and i know that uh pulling together that cfo dashboard i think at one point took days of etl work and i think you've really modernized that and gotten it down to maybe even hours absolutely yeah that was one of the challenges that we had with the cfo dashboard i mean we it took five days to move or etl the data from the fms source to our landing servers and then you know displaying the details for the end users because our end users wanted a lot of details uh on a single sheet that would either time out or it would take way too much time for it to be useful for our end users experience that's really really cool how what is the uh goal at the va i'm gonna kind of move on to the next uh part of this which is like what do you hope to do in terms of improving the lives of veterans or improving the way the va operates through data well for us since we are a financial services center the value that we're providing is giving financial decision makers uh good analytics with like i said with cfo dashboard with supply chain dashboards even um when we have denials with claims understanding why and and even getting to the point where we could be predictive uh and go out and find out how we can stop those denials or at least change and modify the processes so that those denials don't occur as frequently so for us it's mostly making sure that the va operates efficiently so the money could be used well for supplying services to our veterans seems pretty obvious and uh hopefully there's efficiencies that can be brought there and uh you know better benefits for uh veterans that's what we're all about hey andrew maybe you could uh talk you i know we we had a pre-call and you told me a whole bunch of things about how uh data and analytics have improved the lives of uh citizens in hennepin county maybe you could share some of those stories here sure yeah um so i think sort of the
first way um is you know power bi was just adopted a few years ago by our organization it's really taken off um and one of the things that's enabled us to do at least more easily is to be much more transparent so we have a number of public dashboards now that we share via power bi so we're just giving the residents of hennepin county a sort of a full view of okay what's going on uh for example i i own the one for the county attorney's office right so what does that five year look back look like for our case loads what types of cases did we get what was our conviction rate all of that stuff we're putting out publicly for people to see to break down by geographic area including also the racial demographics the gender demographics so people can see okay what sort of disparities are we looking at here um and of course when you do that along with that comes pressure to say what are you doing about those disparities um and that's something we're doing as well you know we're building um we have a disparities line of business we call it so we're really focusing across the board on how do we reduce these disparities now that we've been transparent about what they are where they are how do we set up programs that target them to reduce them across all of our businesses and i think it's important to note that hennepin county is uh the county that uh contains minneapolis yeah and uh you guys were uh recently in the news as a result of the george floyd uh case so there's a lot of racial uh concerns and uh requests for parity there yes it is a very uh tense time that was that was our building on the news every day for a while um so yeah there's there's a lot of pressure and rightfully so to to do more in that area and i think it's a great great service that you guys are providing that transparency and visibility so that people can see where the injustices are and we can resolve them yeah yeah and just one more example you know if i if i may um so another another project we've been working on is um you know so as a county we fund ourselves in part through property taxes and people go delinquent on those every year a certain subset of the population um in a certain subset of those delinquencies you know stay persistently delinquent and go all the way to forfeiture where the county repossesses the property and then sells it off to deal with the debt um and we've been uh using machine learning now to try to predict which of those initial delinquencies that pop up are most likely to make it all the way to forfeiture so that we can then intervene in the process earlier to get those people services help them out with paying their back taxes things like that um and that that's also a disparities project because as you as you may have guessed you know that disproportionately impacts uh poorer people and people of color in our in our county so another example awesome and it's great to see you know tax dollars going to use where we get transparency in that visibility and the county is actually doing things to help the lives of you know those who are less fortunate so kudos to you guys hey paul tell us about this it's cms cms affects everybody in the country at some level i would think uh how's cms using data to improve the lives of uh citizens
citizens well there's a lot of examples uh the fire standard there's a lot of work going on there which is about the interoperability of data sets uh across the country including that comes from cms but i guess the focus that i have is uh you know you really have got to do is the best job you can in firming up the uh you know i'll feel a little chauncey gardner at you from uh peter sellers uh movie being there uh but you know you've got to have a strong trunk and it's in the tree limbs have to be strong you know in order to have the fruit of machine learning and artificial intelligence and some of the other things so what i'm what i'm getting at it is you know what are some new ways that we can provide more complete timely and synchronized data and the data would be better defined and cataloged in other words you know you can bring all the data scientists in and uh and the like but or even supporting a good example would be supporting our constituents whether they're the uh the doctors and the institutions or the uh you know the beneficiaries and the recipients part of that is if i could have all of the information about a claim at the fingertips of the customer service rep or at the uh you know david was talking a little bit about the financials if we're talking about a dispute we may have on a payment side you know having all that information available whether it's a marketplace enrollment or a medicare claim you know it just is going to help reduce the burden another big thing that we're trying to do at cms is reduce the burden on the entire system that our job is to uh see how we can uh you know get the doctors and uh consumers in a better state where they're not overburdened by uh too much regulation or just uh in my my perspective what i work on just uh not well organized data so i think one of our great challenges is uh you know how do you better organize yourself with the tools that we have and uh i'm very interested in a concept uh that we we had a we had a little group that went off about five years ago at cms and we looked at how the cloud might impact us and what we do as cms and we thought that you know organizing how do you organize yourself to really get that data complete timely and synchronized uh and one of the ideas would be is that you know you can't tackle if the enterprise is too big if you're going after the project level it's too small so really the data domain seems to be uh a really goldilocks kind of you know whether it's medicare claims or provider benny those are the kind of domains that if you could organize with what we'll call collective data stewardship this is a concept that we're starting to define a little bit amongst ourselves where the teams that create the data the teams that place the data in an open format they register that data then uh in the data catalog and then that catalog has a well because our data lake that we're thinking about that we're building out at cms is a little bit different than what industry uh has has defined it ours is more of a metadata essence that also has access control wrapped around it so but what that does because our programs are so diverse and the expertise lay at many of our decentralized organizational units that they know our data best and so if we can get the teams that know the data best to define it in a standard data dictionary to then catalog that or place that into a catalog and only our best data we only want the best data in that catalog at least to begin with right so uh that's how at some point you've got to try to figure out like you know what is the level of granularity you're going to organize yourself actually we think it's the data domain okay then what do you got to do just simple stuff the data that creates the team that creates the data defines it places it in a dictionary serves that up into a catalog and then access control on that and i guess that then gets to uh at some point you know i'd like to talk a little bit about uh delta lake uh and and some of the enabling technologies behind the scenes i know you rewrote your kernel i think you're calling it photon um because at some point how do you scale that you got massive amounts of data and different teams are hitting the same sort of uh s3 node to say if you go with amazon um you know is is it told like databricks going to be able to handle the throughput and some of the metadata cataloging that you all do so those are some of the the things that we're exploring you know yeah those are at the end of the day though it's about better data i think that's how you serve the public more complete timely synchronized data you uh you just uh summarized it incredibly well and i was going to try to do that but it's about better data for better uh better decision making right so hey i'd like to move on like
you guys have all uh shown some really great successes but how about some of the challenges that you've run into through this journey and uh you know around data around collaboration maybe around training paul you want to start off since uh you still got the mic oh that's a tough one uh because you know we came from we had a 35-year run with the mainframe mentality and i'm even throwing the mid-tier in there you know like the oracle stuff um but uh you know it was a great run you know people are still using it it still works but man the cloud is offering a lot lot more at a lot reduced price but it's not just the cost it's just the whole ability to think and transform not just analytics i'm even throwing i have a pretty strong background in operations you know we have payroll personnel accounting those are some systems marketplace so i know i know what it's like to build an operational system too and i think that the the line between operations i'm not necessarily talking transactions but the line between operations analytics can easily blur you know if if i've got the operational team uh laying their data in some order of operation and it's out there in an open format why can't i also have analytics sit on top of that now that's something we know we can do but how do you how do you have that conversation with your business leaders or even your i.t leaders that have all come from this mainframe concept again it worked really well for 35 years i mean cms does a really good job of managing its data for the program but we're now sort of in this point where you know the texas power grid worked great until it didn't state unemployment uh systems worked great until they didn't uh and so uh you don't know what's coming next you know as a nation i think we've all got to shrink that security perimeter a little bit by bringing data copies into maybe a more unified uh environment so i guess the challenge is it's one of indus it's not just cms it's like i think it's going to take an entire generation to figure this out but how do you move on i mean so yeah i was i used to have a phone i picked up and i dialed the rotary dial then the flip phone came out you know then you know blackberry now i have an iphone 11 you know and i'm soon gonna get something on my my wrist but that's how does an organization or you know how does an organization make that turn and that's where i'm coming back to the uh thoughtworks uh data match by uh that's because i think you've got to figure out how to organize both operations and analytics but uh so i think it's a culture i think that the challenge is by far cultural and i don't mean that in a derogatory way because we're coming from a really strong position but yet there's a new generation that's got to take ownership of this and take it to a different level because uh we don't need to hoard the data like we do today there's so how do you organize yourself and i think technology is the least of it uh but it also is the enabler so i guess that's uh boy if you could if i'm like you're you're the smart advanced consultant you tell me how to do it because you know you talk about hoarding data um i was just reading an article this morning about you know people talk about data as currency and i'll give you my data but what are you going to give me in return and data you know we talk about democratizing data but when people treat it as currency there's no way to democratize it and people think they own the data the dod again in this report that i was just reading today said hey the data's not your data the data's the dod's data you don't get to decide and i think that's probably true at cms and everywhere else in in the yeah let me throw another shout out there to the department of defense they put together a really nice set of principles that uh that i was able to look at they're real good now we're working tweaking those ourselves within cms to make it a little bit more cmo centric but uh they had a principle in there called the share first ethos uh go from an organization that hoarded data again i think partly in in uh in light of the mainframe limitations but uh how do you actually uh share that data and i just think tools like uh spark uh that's part of the answer but uh yeah i i did want to give a shout out here department defense has done some less work in this area yeah hey dave uh you must have some points on this yeah so paul you mentioned how we organize ourself i was thinking about how you know where do you put analytics from a large organization's perspective do you put it in the business lines do you put it in i.t do you put it separate from both of those and i think all three are the answer uh you know you you need to have it in ite and monitoring it's analytics uh so that you can perform well but you also need it in the business lines like we have in vha with informatics you know the client clinicians working um the analytics but uh and we have it in a pseudo business line called finance you know but it is a business line uh but i think you also need it separate from all of those so you can look at the whole organization and serve like the secretary so i think organizationally figuring out how you centralize and decentralize analytics is a key uh element in this especially for the sharing of data that you're talking about these data silos if you have that centralized control to make sure that the sharing is occurring but also the decentralized capabilities uh serving their masters i think it's it's important way to go but one of the challenges that you mentioned was um uh you know capacity on that on-prem capacity we have machine learning uh that used to take us 24 to 48 hours just to run this machine learning and now that we've gone to this cloud environment where we have the data bricks uh um churning that data and uh we're able to get it down to two hours from 24 to 48 running these complex machine learning algorithm so um there's there's a lot of uh resolutions that can be happening with these these tools but one of the challenges for all of these tools is that there's so dang many of them it's like walking into the ice cream store and you got too many choices and and adding to that it's like they're continuously bringing out new ice cream choices and you know we got the choices for the cloud platform you got the choices for the uh the databases you got the choice of data management for the visualizations and the whole devops pipeline there's so many choices to work with and they're all improving uh exponentially the exponential innovation curve is is making it difficult to architect well and stick to it because the challenge of not sticking to it is a whole bunch of tech debt every time you modify and change to a different platform it takes a long time to get that back up and running and operational so those are some of the challenges we face [Laughter] we talk a lot about the fact that people have you know a separate set of tools for you know data engineering and data transformation then they have a separate set of tools for data warehousing and maybe another set of tools for machine learning and if you've got real time ingest that's yet another set of tools and you know there's there's firewalls between each of those so the whole collaborative capabilities that you know you can have a single set of tools that uh run through all of those different personas and use cases uh can be uh very liberating yeah and i know you guys use databricks for etl for some data warehousing as well as for the machine learning so you get you know a really clean path there exactly hey andrew what about uh hannah pink county what are some of the challenges that you guys have and tell me uh how much of them around skills jared well certainly that's one that's gonna say i was gonna echo what paul and david said um certainly that's something i'd like to hear paul's thoughts on actually sort of this evolution of and david as well this evolution of mpp services to sort of take on the functions of spark is kind of throwing my work um up in the air right now so in terms of what we're going to do and what we're going to use um so that's certainly a challenge the silos as well paul mentioned that um you know people owning the data really sort of i think there's a lot of fear around that for me it's not been technological it's been well we don't want to put the data over here where everyone can get it because then the analysts are going to use it in things we don't know about they're going to misinterpret it they're going to misunderstand it then there's going to be all these problems um you know to me i think that's i think that's sort of a poor way of thinking about it like i think people need to make mistakes they need to learn that's how people learn when they have a lot less fear about that and that sort of feeds into the skills problem too where you know if you hire people and then you lock them into just basically writing reports off of what views the dpa creates for you they never really learn the data systems underneath they don't really know how to query them they don't really learn going back to my grad school background sort of the data generating process underneath how that works and so they're not able to explore ask better questions generate better insights things like that so it's certainly hard yet to hire to find people with the requisite skills but i think we often shoot ourselves in the foot by not by sort of over governing the people we do have go ahead so when we have a data breach or something like that it really increases the pucker factor and so that's uh ends up having a almost sometimes a pendulum swing that goes too far with controls and then then you end up with having a difficult time getting access to data to do the right things so i agree with you andrew that's uh definitely something that we need to work on so there's another uh idea here that came out of department of defense uh it's a principle called data fit for purpose and what that gets to is that yeah i need to define the data via data dictionary and maybe have a community in the catalog that uh that talks about that but really if you can put the business rules and uh we're blessed that the cms to have a pretty robust enterprise architectural program so uh i guess andrew to your point though you know some data really can you can impale yourself with your own ignorance on data and you do need to be careful um about the way what data that you present to others to be consumed you know the database keys might be squirrely you need to you need to write about that but i guess the idea would be if you could get as because we do have these robust data catalogs and they're more than uh you know what you would think like a dewey decimal system in a library these can be whole ecosystems in and of themselves where you can put links to business process models that your enterprise architects let's say i have a data domain i can model the major uh data aspects of excuse me business process aspects order of operation who's going what the citations to the requirements etc uh you could do the same thing for your data most data domains have more than one data system so you can start to show an assistant's interaction model uh how those that data flows through the system so i guess when i'm going the data fit for purpose if you're really smart about it and use these tools in a synergistic fashion you can have a broader context for where this data came from its creation story the restrictions there about it because at some point you know you can empower yourself on your own ignorance and there's got to be some sort of uh organizational response to controlling the access to that while you're also balancing hey you know people need to as you say they need to learn and to grow and understand the data so i think the challenge here is to engage a whole of sort of government kind of approach to bring more than one type of uh context to the data but in a very uh wiki kind of easily understood linked way uh to give that data fit for purpose is this the data that i need to use and if so tell me everything there is about this data so that i can interpret it correctly and that's i think that's yeah that's one of the answers to that hey i'd like to um move along a little bit bit and uh ask you and dave maybe you could start off here with um you've talked a little bit about the role that databricks is playing in your environment but maybe you can tell us a little bit about how are you using databricks today and uh i guess i'll also ask how important is openness to the va and to the work that you're doing and how important was the role of openness and data bricks in your selection process there i probably threw a bunch of questions at you there didn't i yeah that's all right so pick any one of them well i'll start with one that you didn't ask it's a little less obvious um and that's the positive impact that databricks had through the customer service i mean uh we were put together with a dynamic dual with both databricks and microsoft working together to make sure that pipeline is working well for us and they i mean they consistently support us on the scrum calls they were persistent to resolve issues challenges convic they made sure that we uh used configuration best practices and they helped us enhance the overall performance so that's one aspect that i want to you know that was very helpful for us because uh when you're starting and putting in that tech that new uh technology into your environment uh there's a lot of questions that you know you could really screw up and and pay a lot of money for without knowing it and having those resources on board to help us out was it was fantastic but you know we already talked about the performance issues from five days to a couple hours um you know this is i talked about the size uh 10 billion rows of data in the fms database when we're talking about eight years of data now before on-prem we only could do two years of data because we just didn't have that capability and putting this in was allowing us to you know enhance the user experience by having much more uh research capability into that uh eight years of data versus just two uh and so yeah there's also this other aspect that really has made a big difference and that is um when we're on prem there's a number of tools that we're coding in a number of um areas that and we're having to chunk up this data so we have complexity in the code and everything and then once we've moved into the databricks environment are data science where it was able to just really reduce the complexity of the code work in one environment really make that etl much more smooth and uh efficient so that also helps us in a huge way and that's maintenance we don't have somebody have to scroll through there and figure out what's going on uh with thousands of lines of code now we've got some efficient code that uh it's it's very easy for other people to come alongside and understand what's going on so it's it's benefited us in a lot of ways as far as like um the performance the uh the folks that's come on site come alongside of us and also the complexity so we'll definitely appreciate that feedback and uh i know some of the people that have worked with the va directly from microsoft and from obviously the databricks team and uh they're top-notch so glad you appreciate that hey andrew tell us a little bit i know you shared a little bit about the fact that you know python and arcane plugged in what are there other aspects of uh uh you know databricks in your environment that have benefited hennepin county benefited hennepin county oh yeah well i mean certainly the automation has been very helpful for us and i just think some of the things we're talking about here about just being able to start to use lake architectures which you're still relatively new at to sort of open things up so that people can can sort of get in there and see what's going on in the data and databricks makes that that pretty simple to do particularly if you work with the hive tables um it's very nice you can kind of get get a big structure going it's almost like a database um yet you know maybe for whatever reason is governed a little differently um and so people can get in there and see and if you know a little r a little python even just a little sql you can start to examine what's going on um with that data in your lake and start to learn start to learn new systems um which has been really great so and andrew are you guys using kind of the lake house construct which allows you to put the data in the data lake and then bring your sql bi tools directly to it yes that's what i'm pursuing um with the public safety data lake yeah which is sort of our biggest data lake project going right now yeah i would like to keep stuff in hive as much as possible certainly there's it's interesting there's sort of increased pressure to move that along to a database to a regular database and i'm resisting it we'll see how successful i am um just because i this gets back to what we were talking about earlier with governance i'm sort of afraid um that as we move to a database or as mpp starts to mimic those spark functions and it so maybe starts to take over some of the late processes that we just get the same old governance structure applied to that again and then i think people are are limited again um and i really don't want that so yeah so we're going to do using that we're using delta lake and it's been it works pretty well i would say so far you know i'm certainly looking forward to photon and sql analytics i think that'll up our game another notch so yeah i think so too and uh just going back to this dod thing again the dod specifically said in a memorandum that they want to store data in a mem in a manner that is platform and environment agnostic uncoupled from hardware and software dependencies that's kind of the the vision of lake house once you put it into an mpp you're probably putting it into their proprietary format does that um does that uh weigh in your opinions yeah if you're asking me yes you know to the extent i have a say you know i'm the data scientist so i don't get to make the architecture decisions um so yeah i mean i worked with microsoft's tool recently just sort of trying it out and uh you know it's it's in preview still so they're still adding functionality to it but just some of the capabilities databricks gives you in working with delta are certainly at this point you know far superior and much easier to use um you know stand out for you andrew yeah uh well let me think of how to put this this well i mean it just wasn't possible to read the delta files natively um in the way that databricks does it wasn't possible to work with the log data so you have to do set up a work around and i don't know if that will be fixed in the future but um i thought that was pretty limiting so paul what about you what's the role databricks has played at cms i know we're involved in i don't know seven or eight different programs up there any that you're directly involved in well yeah i've seen uh seeing a bunch of them come them come uh on board you know you got the uh the work that uh started in the medicaid uh area uh with some very pioneering work in the sense that uh yeah brought a new tool like it's a new class as well i mean databricks is just a new class of tool uh but the thing i'm keeping an eye on now is uh you know it's still not in production yet but the the fraud prevention system our uh uh center for program integrity has done some really really great work uh sort of taking the initiative on a lot of the uh machine learning so i think and artificial intelligence so they're so sort of in in a way you know you're trying to move as much of these new technologies up as far in the life cycle as you can and then of course you have uh activity that's going as data's in flight you could interrogate it there and then once it's sitting in a big pool of data for a very long time uh you can do some more uh you know different types of uh deep learning so i think all of those are happening at cms um and so it's it's very exciting i i do think the open isn't that different too right isn't like the like the whole open source started with the coding frontier and the whole agile movement and i think you know and then you know so look look at what happened you had the the main the hardware folks did their thing and then the software folks and i think the data people are just bringing that third wave of change and uh how does all that come together i still think we're sorting it out but uh i really just the notion that not only is there open source code for data tools like spark and that you guys donated uh delta lake to that which is a very cool thing but then the the the open standard itself you know like say like a parkade uh it just opens up a lot and i think that's hugely important to see us i think it's uh it's a discriminating factor actually you know it's a it's a very big deal because it allows us to if we if we want to pick up shop and go to the next cloud that has a better offering or more secure offering you know we'll we'll be able to do that and i think that's you know you're always going to be locked into a certain degree but uh i think i think this whole movement of open the open movement is a very big deal and um taking a look at any of these lake house architectures that dave bricks and others are promoting uh i haven't in about a year you know so i try to i've been going head down trying to get uh some data strategy work but uh i do uh david was talking about the pace of uh change uh i had the privilege of helping to organize that our first data summit several years ago and we brought a bunch of technical folks together and we did it out of necessity because we were just getting into that i mean you spend your whole career learning about db2 or oracle and you're you know fat dumb and happy and then the next thing you know a bucket of cold water gets thrown on you and you got to understand like you know what's s3 you know and the list never stops and as david says it's accelerating uh so that makes it uh quite a challenge to to stay up uh you know with what's going on but uh yeah it's it's interesting where this is going and it's i i think we are hamsters on the cage david i hate to tell you i i think it's we're on this journey and it's not gonna it's not gonna ease up so well that's a great segue into my
final question for the day like and that's kind of where do you go now in your analytics journey and you know what's next on the table for a cms call wow um well i think there's just a lot of of again i've been at it for 22 years at cms i've never been more positive on where the future can go there's so much innovation that's taking place and i think there is a recognition that we've got to sort of uh have some governance we're not quite sure what that is i think the whole industry is you know it's a lot easier to talk about code and business requirements data has its own world and you have to sort of learn it and so i think data governance is going to be a big deal i think new architectures like data mesh uh is a big idea you know that the ghani's papers are good uh and then i feel a challenge uh to folks so and uh howard you turned me onto this few weeks ago so okay so now there's this new thing called feature stores and i spent some time with a good thing about the internet is you can go i i went and looked at some of the past sparks so much there's some fantastic uh stuff out on the web that gets you up to speed fairly quick but you know if we don't watch it if we don't get pre-organized like we know that in a couple days all these cicadas are going to come out of the ground right and and we can sit there and watch all this come out and and uh but how do you there's a guy that's put together a tracking mechanism for us to become like a scientist to help him figure out where all this brood x is going to come so if there's a you know we get tens of thousands of people taking pictures of the bugs you know we can then uh collectively uh get a better sense of where we're going so i worry the same thing for like machine learning and wouldn't it be nice if uh if we had the foresight to organize ourselves more now to get this metadata under control for the things that we already have and then keep an eye on what's coming next as we get these big libraries of machine learning these tools are coming they're going to have their own uh sort of metadata that has to be integrated with this bigger picture so i guess that's what's coming next and i i think the next generation i think we've made it through the cloud we made it through the hardware we made it through the tools like spark i think the next challenge is uh metadata management and knowledge management and bringing all this together so that people can find what they're looking for more quickly they have trust in it when they get there and then we let a next generation of people come through uh innovate i think that's you know but but to do that we should be laying down more organized just get ourselves a little bit more organized from our governance and i think it comes back to metadata i think you're right paul it's uh the next frontier you know not just the metadata but the governance of the data i think all of that's uh you know connected together and uh it's still a challenge hey um you're not gonna you're not gonna talk about the cicadas though i'm still waiting they keep saying the cicadas are coming any day but i still haven't seen them what's the cicada no i'm just kidding so similar to how you uh talk about crowdsourcing the cicada pictures um you know i think that one of the ways that we want to move forward is like with hackathons because now you bring in a bunch of different folks that are doing it differently in multiple locations together to learn together to grow together to understand what's going on elsewhere uh start learning about other tools that might be doing well for other folks um and so that's one of the things that va is working on is putting together some hackathons and we're even wanting to host that to bring together a bunch of folks to take use cases and explore them with the big data that we have available from the fms data to other types of data that we have available and even databricks is helping us out with this one to um to work that hackathon together but we want to we want to start that's that way of getting that knowledge management kind of centralized and working together but decentralized and so i i'd also say that it's not just that collaboration but one of the things that we want to do is make sure that we paul you talked about architecting and being able to move and being um so you have a conflict between being able to quickly change cloud platforms or whatever and the tech debt that that increases so as new technology comes along reduces that tech deck so you can be uh more versatile i think we need to be careful though on how quickly you move you need to be somewhat decisive and yet being able to move you know so there's some deep challenges to to those decisions that you make early on on what architecture you're using and that's what i like about databricks is it's one of those that can you can move with so uh yeah that's open source type of concepts is going to be very important in the future yeah and uh thank you dave and andrew um what's uh what's next for hennepin county yeah so i think the next big step for us is you know we're still finishing this public safety data lake project and now we're we're talking about scaling it so can we do something like this enterprise-wide where we get all of the departments and lines of business together to create sort of a single source we can hopefully develop a single identifier so we could see things like you know if somebody was in the jail were they also on you know food stamp service snap service um where else have we been touching you know these residents in our system and can we use that to our advantage to help them out um so doing that figuring out the platform for doing that um is that certainly the next step um after that i think maybe this will go along with it is i think doing more with our unstructured data so everything so much about relational data what are we going to do with it where are we going to put it how are we going to provision access but there's so much unstructured data and so much more coming in day by day just in public safety you know where body cams were added in the last couple of years um there's transcriptions of so much i mean now we're auto transcribing certain things um we just have lots and lots of digital data that we're just not doing much with other than storing and certainly there's a lot to explore there so yeah that's a threshold for everybody to overcome and many organizations are doing much more with that today but it's still uh definitely a challenge and yeah hey um let me ask uh one final question um what lessons have you learned andrew can you share any lessons learned at hennepin county something that somebody could leave here and say oh boy i'm not going to make that error or uh i know the path forward based on what andrew had to experience learn from me um well you've been down the path a little bit so uh you might have some experiences yeah um that's a good question i think you want me to come back to you yeah maybe circle back in a second yeah let's think about that okay dave you gotta answer for that one sure i think for us for me um personally learning that really you need to go back to the business and understand what their outcomes are gonna be needed you know we can easily get into some data science and go oh i've got an idea and we can go all over the place on crunching data and coming with answers but what is the outcome that the business actually needs and we need to keep going back to that and and deeply understanding that to provide them good products that will will support their needs and also a lot of times the business doesn't know um what they need what can be done uh and so providing them uh concepts and capabilities that can actually accomplish like i think paul you were talking maybe about incorporating the analytics in the software systems so that the software system doesn't need the human to uh engage in certain decisions you know if if we have good uh analytics in in certain decision points so even knowing that that can be done is a big benefit to the customer but a lot of times they don't think about that so getting the uh getting into the the customer and the businesses um business is going to really help us useful come up with useful products yeah what's the art of the possible and making sure that it's aligned with the business value i think those are great points i don't know if paul you can add anything to that that kind of summarizes a lot for me yeah it did for me too i think that's the trick um you know i t can be an enabler but you've got to just be patient with each other the business and i t have to be patient but they have to talk uh and they have to keep working through ideas because there's a there's a change of foot but i agree it's uh the business is a you know that's the key it's it's really the key to solving all this being with the business but but also what you're saying david let them know it's possible it's sort of a two hands on the steering wheel yeah andrew right back to you for final comment yeah um i guess when i think about some of the challenges um in building out something like a lake architecture in your organization it's it's how much relationships matter um you know at least i think it's sort of probably different in in private sector where my friends tell me you kind of just have to convince the cio or the cdo that you want to do something and if they buy in then you're good you know that's the way it feels talking to them but in our organization the structure is just a lot more flat and so you can't really you know just convince somebody who's in administration and then everybody buys in and they do it no you have to leverage those relationships you have to have those relationships to convince people and to build trust that okay we're going to do this and i'm not going to cut you out i'm not going to take over your work you're going to be a part of this you can use it and leverage it to do better work and sort of going along that journey at the pace that the people around you will let you go i think is important i love that and uh i always tell people we're all in sales we're all selling an idea and trying to get people to go along with us so uh i think that's great hey i want to uh conclude the panel uh today i think there were some really insightful comments from all of our panelists david andrew paul i want to thank you guys for your partnership uh for your uh support and for uh participating in this broadcast today and uh thank you for your uh you know using data brands so thank you thank you guys what a great discussion so many best practices and lessons learned thank you to our amazing group of panelists and thank you all for tuning in please make sure to stop by the solution theater later today for interactive demos on the most popular data and machine learning use cases across industry enjoy the rest of summit [Music] everyone everyone [Music] [Music] you
Public dashboards reveal disparities.
Strategic advantage through data.
NoSQL to Lakehouse success.
Culture change for data.
Billions in savings possible!
Trust and collaboration are key.
Lab vs. real-world AI.














