In an era of unprecedented data generation and complex societal challenges, public sector agencies face a critical dilemma: how to leverage vast datasets for public good while rigorously protecting sensitive information. This session, featuring Dave Thomas from Deloitte and Danny Holloway from Immuta, dives into the intricacies of implementing effective data security and policy frameworks, transforming traditional 'need to know' paradigms into dynamic 'need to share' capabilities.
“Data has become democratized and the tools have lowered that barrier to the point where we now have more actors that that have problems that they can address with the right tech technologies and the right rules around the information they have.”
Data sharing in the public sector is critical for addressing complex problems, yet fraught with security risks. Discover how to balance the imperative to share data with the need for robust governance, transforming 'need to know' into 'need to share' securely and at scale.
hey good afternoon everyone thank you so much for joining us this late in the evening it's my pleasure to introduce you to dave thomas from deloitte and danny holloway from immuta who are going to be speaking to us about security and governance we always say security should be at the forefront of a project but here we scheduled it at the end of the day so i don't know somebody didn't get the memo so in any case welcome guys thank you so much thanks appreciate it and thanks thanks for coming uh really excited to be here as we were reflecting on our our time slot it reminded me of the upcoming fourth of july weekend this is the last thing that i have between that and and celebrating and i was remembering as a kid going to see the fourth of july fireworks and every time a couple fireworks went off right next to each other i'd lean into my mom and say is it here yet is it here yet and finally they'd all go off at once and it was the grand finale and so here you guys are you made it it's the grand finale of the the conference so really appreciate you spending it with us here today and danny if you want to introduce yourself yeah absolutely thanks dave so i'm danny holloway i'm the cto for public sector at amuda we are working on automating secure access to data and really excited to talk about some of the lessons learned and cultural practices that we've adopted in rolling out data security frameworks for our federal customers and i'm dave thomas i'm a principal in deloitte i'm one of the leaders in our ai and data practice and i focus primarily on data ops and so when we say data ops what is what does that really mean it means driving value driving a data value chain incrementally but in continually with our clients and one of the best ways we can drive that value is by sharing the data so that's what we'll talk a little bit about here today how do we how do we drive that value by sharing the data and i work across all of our our sectors so civilian health health defense security justice state local and so the the lessons that we've taken away the from implementing these solutions really apply across the board uh so so we'll sort of talk about a number of the the challenges we've seen but you can really apply it apply it anywhere
and so we will talk a little bit more about some of the challenges um and some of the the tactical things that agencies are weighing uh when they're when they're thinking about how to share data uh and then we'll we'll talk about some of the reasons why sherry data is so important and then danny's going to talk a little bit about some of the ways we can we can make it easier because it's not a it's not an easy thing and then finally we'll we'll chat a little bit about what when you go through that what are the outcomes you enable
so so first why why share data and i was on a run this morning um and i love running in san francisco every time i come here i go for iran but it struck me as i was running that i had heard the news about increased homelessness in the city and it was it was really real as i went on my run this morning and as i was running i was thinking you know as a public sector guy well this is a public public sector problem how do we solve this problem and i began to thinking about the data what are the types of analytics we'd want to pull together and it struck me that the data comes from from all sorts of different agencies right so there's obviously security concerns there's health and human services data there's economic data housing data mental health mental health data and in order to to really make a difference in in this in this challenge agencies from across the spectrum really need to share that data and then continuing to think about it though the data we're talking about is unbelievably sensitive so you're talking about uh you know health data and personally identifiable information and law enforcement data in some cases employment data or or housing data and so to to get all that data together you really have to have to respect all of the different stakeholders in that so the the individuals themselves the the people who are sharing the data um you want to make sure that your data that's been collected for for one use may be appropriate in certain situations but but not appropriate to share another so health data is a great example right and so uh so that's just an example of all the different places where where missions are continuing to expand and crossing would have been organizational boundaries and really require agencies to come together to share data in order to make some of these differences i think the other thing that is that's
really struck me over the the past couple of days here at the conference is the changes in the technical landscape as well so as we look at um uh the new technologies when sharing data when sharing data even just a few years ago it would mean attaching a spreadsheet to an email and sending it to somebody or or maybe if you're really fancy you'd have an ftp site and you uploaded it with a with a cron job or something like that and it just doesn't doesn't work anymore um the the scale with which we share data technologies like you know delta sharing or um these agencies are building these these massive data lakes these massive data warehouses the lake houses etc and the way people are consuming data has changed as well the analysts don't want a spreadsheet that's already been processed they want the raw data if you're doing machine learning on the data you you don't want it cleaned you want to get to the raw and you want to build your own features and i thought we thought about a number of the ways we used to to talk about data sharing in um in dod and the intelligence community and it was always need to know right so someone asked for data if you didn't know that they had a need to know you'd go and ask somebody and it was human adjudicated need to know and and that was the speed at which we we were exchanging data and and it just doesn't work anymore the scale is too big the the types of data we need is too big and the speed with which we need to share it is too big and it's not always even person to person often it's it's machine to machine and so how do you implement policies that enable us to enforce it regardless of how data is being consumed and not just not just handed over through an email or a spreadsheet
and of course the consequences of of getting it wrong are very real they're very tangible and and people understand those really well um and and sometimes the the benefit of of sharing data can be a little bit ephemeral right so especially when you're thinking about data analytics and data data science right data science hypothesis testing you know if we bring this data together we might get some insight that can make a difference in this mission so you're you're constantly weighing the hypothetical benefits against the real and intangible costs of sharing data right so data could could get out of out of out of control someone could share it again to someone who's not supposed to maybe you've crossed an authority a boundary of authorities and you're sharing data to a a part of an agency that shouldn't be able to see it and so there's all these real or in in pii getting out personal health data getting out and and the the trust that all the citizens have in in government is a real issue as well so if they think data has been shared for for one reason and it gets used for something else uh the erosion of trust can't be can't be rebuilt and and so we really need to uh pay attention uh to that and this is actually an example where i think um you
know the the federal government has been ahead of of the private sector for a while uh starting to see things like gdpr or the california consumer privacy act i noticed here when i when i flew into california i got a different notice on on pop-ups on my phone i hadn't i guess i hadn't been since that log out passed so i had to figure out what what were my answers which i usually don't share anything uh and it was amazon i think just a little bit ago had to announce a 750 million dollar uh fine for for inappropriately using data data they've been collected in one context and used in another and we're we're very comfortable we've been operating that space in the public sector in a long time so it's uh it's kind of interesting to to see the the the private sector catch up but again it's that it's this balance where we're we're constantly um uh operating in what is the what is the cost response what is the cost and what is the benefit and how do we start to to even the scales a little bit what are the ways that we can make the benefit easier to get to and reduce the cost or the potential risk of sharing that data the good news is that um there are a
couple of things that are being done to help push that scale so so one of them is an emphasis around the value of data and around the value of sharing data and so on the slide here you can see an extract from the federal data strategy framework and the same thing is true in a lot of a lot of agencies there's a real emphasis on using data to make decisions so sharing data bringing it together and of course there's the ten principles three of them are on you know using the data and then seven then are there on the privacy and protection of the data so the responsible data but the good news is people are starting to you know push up the the imperative to share a little bit more a little bit more explicitly and so so so that part is coming up and we're starting to see that that more and more and uh the challenge though is that in the in the public sector uh we've been paying attention to the risks for a long time we've been really very careful on the risk side so there's all sorts of uh controls in place and appropriately so uh to ensure that data is not shared inappropriately so you know everything everything we do
in in the federal space and uh really everywhere is uh you're governed by the the privacy act right so what is the data what is being collected for why are you using it in transparency and then privacy impact assessments is a great example of how the government is explicitly thinking about that balance for a given system what is the data you're collecting what is the data you're using and what are the benefits for it and do those benefits outweigh the the risks and civil liberties and um um civil rights lawyers are looking at that constantly to decide whether or not that balance is is right and even when you get to get through that and you start to talk about sharing data there's all sorts of explicit constraints around sharing data right so interagency data use agreements inter memorandums of agreement or understanding that explicitly codify what data can be shared for what purposes what under what authority and to whom and these are always point-to-point agreements or typically to point and then the uh the authorities to operate right so now now that we've gotten through and all now it's allowed to be shared we've made sure we've met all the policies what explicitly are we going to do in our system what are the ingress and egress points for the data let's explicitly say that and make sure that those are very well documented and cleared and what data comes in and out and how we protect that that data another example i i thought the uh the the introduction of the data marketplace from databricks was really interesting uh and again a place where i think the public sector has really been pushing forward but with challenges in that a lot of the data um you know even in the in the pandemic obviously a lot of the data was just made public for everybody vaccination rates infection rates things like that but a lot of that data has to be shared point to point has to have very explicit authorizations and can only be shared with subsets so some very unique use cases in the public space but again even though we have all these these the you know deliberate effort to bound risk to and control risks we are starting to see people again try to push up the value of sharing data and the foundations of evidence based policy making is a great example of that and it really does two things in regards to to that first it explicitly um uh commands that people will use data to make uh a policy decision so uh that it defines the the data that needs to be collected and the data that in the uses and then also it it employs agencies to start to modernize their data infrastructure start to make the changes um that allow them to start to discover um discover the uh the the uses and and uh and make the data sharing even even easier um so um so with that um you know again really this this thinking about this this trade-off the the hypothetical benefits and the the potential risks and costs to getting that i'm going to turn it over to danny and he's going to talk a little bit about some of the the technical implementation uh concerns some of the requirements um that we need to think about as we're doing that and some of the solutions that are available today to get through some of this this balance and and sort of right the ship a little bit thanks dave thanks dave so dave talked about the
the need to know and i'm going to talk about sort of this transition to a need to share and a need to scale and beyond that the need to be successful you know we've talked about issues that are critical to each of us in this room things as simple as national security no brainer but we're talking about things that protect the way that we communicate the way that we commute to work the way that we protect ourselves from current and future pandemics uh the government has a pretty critical mission and so just taking into totality the span of influence and effect that this organization has and the assets under control and the people that they're working with and some of the unique propositions that we can approach from a data analytics standpoint think about just the veterans affairs for instance you know in a single organization that has medical records for people from day of enlistment until the day that they die like that is a unique data set that may not exist anywhere else um and so you know with that in mind i'm going to talk about some recommendations
and foundations for success things that we've seen when correctly applied and addressed and things that are materializing in some ways today and some of the gaps and how to close them so you know first and foremost you know kind of foundation you know you got to build a house on a stable foundation we need a data platform with adequate storage and compute resources it seems kind of funny to say that out loud and need to put it into words on a slide but we have legacy systems that have been in existence for multiple decades at this point and we're trying to approach problems that exceed the scale of what can be processed on commodity hardware and so with the power of of the cloud and multiple clouds at this point we've got some some new techniques and approaches and capabilities that we can bring to bear against these big problems and so that's sort of the starting point you know table stakes uh then we've got to have a process for establishing establishing governance and i can't emphasize this enough it's you know i'm a cto there's a technical talk but this is a cultural thing this is you know an actual giving someone the designation and authority to make a decision on what can be shared under which circumstances and why and further to take that beyond what historically has been this binary decision of yes or no and get to more nuanced engagements and ways of evaluating information and saying i can share this with these people for these reasons i can share this if i redact these values for this purpose and this reason for this amount of time or if i put enough anonymization and protection of the individuals in this data set in place i can start to use it for new purposes that were never even imagined or addressable and so we get beyond you know what we can do with copies and what we can do with these you know finite rules and decisions made by one person and you know tracked in an mou that got filed from one you know person emailing another and we make a scalable framework for addressing these solutions and doing it across technologies across data sets and across environments and networks and possibly even across agencies and organizations in and out of federal so the next thing we need and you know if anyone's got ideas on this i know we're all hiring in this room but a workforce with the appropriate skills and authority some of the ways that we address that are by lowing the burden the bar you know how do we make these you know how do i allow somebody to come in and stand up a spark cluster without having to know the ins and outs of every component that goes into the open source architecture that goes behind it luckily people like databricks have taken that mission on but how do we bring these citizen data scientists and this data together in an environment that allows them to safely experiment to do things with information that previously may not have been possible or may not have been you know cost uh sensible uh and do that in a way where they're not worried about the ramifications of leaking information or you know potentially creating a spill uh and then you know finally this kind of down to brass tacks and tactics but you know an authoritative place for the inputs that we need to make these decisions so we need to know what are the policies that we are you know underneath and need to operate in accordance with who am i getting information to and from and how do i know you know what capacity they're acting in just the dynamic in environment that we're operating people change jobs people change missions they might just you know move from one project to a next and if we're thinking about things in this attribute based dynamic way their access should change when those roles and responsibilities change and often times we see when that's implemented in a manual stag stagnant static process those changes don't happen i've gone into places and been on a list of names that i should no longer be on and then finally metadata how do we do this in a scalable way how do we do things where where we're no longer requiring human curation a manual review of data sources where data's effectively being sidelined as it's waiting to be reviewed and adjudicated and how do we start to do things like classification with machine learning and pulling information from current existing authoritative external catalogs to be able to make these policy and protection decisions has data is coming into our system so that we can onboard new data sets bring it to the fight and use it as quickly as possible
so here's a couple of trends that we're starting to see and you know i'll talk about this in the lens of technology but i think it applies across a couple of different things so you know you know very concretely there's been this massive shift towards dissent distributed computing the flexibility it provides the cloud you know this is kind of obvious we've moved from scaling up buying more expensive exquisite hardware or scaling out rather than scaling up you know two leveraging multiple resources spinning up spinning down doing it in a cost-effective way and what we've seen with this is that more problems are getting approached more more data is coming into the problem or into the problem space to come up with these solutions and you know just like we've seen the separation of compute and storage in the in the you know virtually limitless storage uh solutions that we have from cloud providers we're seeing this need to separate policy from platform so this is something that it's pretty important to mute it but it gets to the fact that we're no longer building siloed applications where decisions about who can do what with information are being made in hard-coded software-defined logic we have applications we have analytics we have exploratory ad hoc queries we have dashboards all being driven by the same underlying information and the farther we can get down in the stack and implement those policy decisions the more ground we can cover the more we can be assured that it's being consistently applied centrally audited and available for that type of analysis and so we think that you know with that you know there's going to be multiple clouds there's going to be multiple platforms we're getting to hyper specialized technologies to address unique problems and if we can separate policy with platform we can have that single policy that's adherent to you know the the law that's driving the title the authority uh the regulation but ensure that it's going to work regardless of where we're processing that data who's using it and why
okay this slide builds so what we're seeing is that these data policies are becoming the roadblock to secure scaling so because we don't know how to operate with confidence and assurance that we're handling information correctly we're either not doing it at all or we're waiting and hoping that somebody else will come in and take the risk for us and so it's not a linear growth it's it's that actual exponential growth we're seeing that with more data sources coming available whether they're commercially available uh whether they're you know being open sourced and collected from open information or whether they're actually derivative data products that have been created out of the fusion of data sources of our current and existing holdings we have more data sources more technologies and that's not going to stop and then we've got more rules dave talked about seeing different pop-ups here in california because of the ccpa gdpr you know with the breaches and the vulnerabilities and the way that mishandling of information is starting to leak into every everyday life i don't see that stopping anytime soon and so you combine that with this explosion of data users we've got new tools we can have you know people who who don't know how to write python who don't know how to you know run a spark job now bringing information into these low code no code environments using bi tools clicking and building dashboards you know taking a dashboard and modifying it for another purpose data has become democratized and the tools have lowered that barrier to the point where we now have more actors that that have problems that they can address with the right tech technologies and the right rules around the information they have so you know at the end of the day we need to make sure we're getting all the information that people should have to the right people at the right time and nothing that they shouldn't have keeps going so i i talked a little bit about uh separation of policy from platform
you know just to dig into that a little bit more one of the things that we're seeing is that the data doesn't live in one place anymore you're no longer going out and getting a license with one database vendor and your enterprise is standardized on that platform you've got teams that are possibly geographically distributed possibly working virtually you know coming in with different skill sets different requirements different types of information and so we don't have to shoehorn information into the decision the the tool that was you know put down from on high by the enterprise we're starting to see people pick and choose based on the skills that they have the problems they're solving and the way that the information lends
itself to being stored and queried and the other thing that we're seeing is that you know these aren't static data sets we're seeing information that's streaming that's coming in faster than ever that's coming from more places you know some of it's iot driven some of it's mobile computing some of it's 5g and you know pervasive access to the internet and some of it just comes from the fact that we're seeing value from information so things that would have otherwise hit the cutting room floor or been ignored or thrown away because it was too expensive we're now holding on to it thinking that there may be unique insights that could be gleaned from that information that might do something as simple as shape our product development roadmap which button is never getting clicked in my system is there a lack of documentation is the functionality behind that button not useful we can start to get these insights that drive every part of our organization and you know in in the case of industry it's typically to drive profits but in the organ the government it could be to operate more efficiently or save more money or save more lives the other thing we're seeing is that you know i talked about this a little bit earlier but you know we can't ensure that it's being consistently enforced some of this policy is being interpreted written in one language and and hopefully they did it right and hopefully they had unit tests to ensure that they did it right and continue to operate right but we all know in reality that's not how it works so if we can define it once and get it at a low level we can put guard rails in place we can put continuous monitoring continuing testing against it to ensure that it continues to operate as intended and we can have it reviewed frequently enough and by people who aren't experts you don't need to know every programming programming language under the sun in order to do a comprehensive audit you can read these policies as if they were written in plain english and then finally um back to that dynamic dynamicism you know people don't stay in one job for their their careers you know particularly in the federal government i think it's one of the unique things about being a federal employee is is you get exposed to different types of missions you know you move across you know directorates or parts of the organization you do a rotation and so there's this really you know good environment for getting a breadth of experience but with that the information you need and the information you should have changes and so you know in the worst case somebody's getting access to information they should no longer have and they misuse it for some reason in the best case somebody shows up for the job on day one and instead of having to go through the checklist and find all the people that they need to talk to and figure out all the information they need to be successful they show up and their data's in their workspace and they can say you know whether show databases show tables and it's everything that they need to be successful in that role today and it might be different tomorrow but it's being done automatically uh and so some of this all builds into what you know nist defined as attribute-based access control and what this helps us get to is addressing a future state where technologies are going to continue to be rapidly evolving they're going to come in they're going to come out they're going to change users are going to come in and go and change roles and responsibilities and without their accesses need to adapt and then data is only going to continue to increase new data sources more information denser data sources and you know i think we're going to see more laws and different ads in different places but without this this system to enforce policy we're just constantly going to be playing catch-up and we're going to either be accepting risk and potentially misusing information or we're going to be missing out on opportunities that could help us with whatever it is that our mission uh requires so with that i'm going to hand it back to dave so we've had the opportunity to work
together and implement uh this type of approach in a number of different agencies and when we do it what we see at the end of the day is that we've enabled uh much more effective interim interagency data sharing so somebody comes in they have access to the data they need on day one data is going to move outside of an agency a boundary that the agency trusts that the the policies are enforced because they can be discovered they can be audited and the data can be seen and they know that all the data policies are stored in one place another great benefit is the speed to add new data so a new data set comes in it needs to be shared if the policies uh are have already been taken from the the legalese and put into you know schemas then you can really quickly enforce that same policy on the new the new data you say this is the same type of data it has to fit this policy we're allowed to share it under these circumstances and you can go off the other the other incredibly important piece is clarity uh about who has data access to what data so our our clients care not only that uh the data isn't being shared inappropriately but they want to make sure they want to be able to audit and they can go back and see who has who has access what data and and what type of of of individuals have access to what data what are the the actual not just the names of the people but why does this person have access what are their attributes and what is the data attribute that allows them to have access and the last thing that i think is really important there at the end is danny described there's a huge proliferation of data stores um the data stores are being stood up for multiple different uh different use cases data is moving from place to place and even today you know in this session one of the major themes about lake house was the need for different types of of data stores right you have your business intelligence and as you move up as they uh they talked about repeatedly on the ai maturity curve you need you need different types of data accesses and that was one look and we see that uh increasingly uh you think about transactional data as well um you do get direct access to the transactional stores because that's the the fastest place to get access so by pulling out policies separating it from persistence it allows you to pick the right persistence engine as long as you can enforce the policy there so that um i don't know if there's any questions be happy to happy to answer any that that might come in hey danny and dave um i had a question pertaining to what you feel the balance is between those like applications and persistence engines like data bricks and uh you know specific data governance software like commuta to do the data governance because what we've seen at this conference is like well as technologies like computer that's meant specifically to do all the data governance uh there's also more maturing on like that the database application side to do that right so what is like the balance between them should it be like a complete overlap with data governance or is it more like a venn diagram where at a macro level one software does this at a micro level another software does the other side of the data governance yeah that's a great question uh so i think there's there's a trade space i would never go and recommend that we immediately start refactoring every application that's out there but i think as we start to build new tools and we start to look at the architectural components that are in our system today that we know about and the tools that are available for us we can start to think about things differently and take a more scalable long-term up look at where we want to put that logic do we want to write it in java do we want to write in javascript do we write in go do we want to run mnc or python or yeah you name it whatever's next or do we want to come out with some sort of external service that can become a decision service that's interoperable you know maybe it's a set of microservices maybe it's a set of embeddable components that can be a part of your application and think about you know how we can ensure that consistency and the time it's going to take off of the people who are actually implementing and validating and assessing that policy and say you know is that worth an architectural refresh like is there a new paradigm in application and analytic development and do we continue to see pervasive adoption of data-driven technologies where where the benefits are worth you know taking a new look at how we make those decisions in any of your efforts have you run into cases where cases where the compliance checkers the compliance checkers are so risk-averse that even when you put a simple straightforward system in their hands you can't get them to implement in it because of the risk of com of committing to a specific way of implementing the compliance guidance the compliance guidance i've got some thoughts you want to start so the short answer is yes um we've definitely run into into cases like that um generally speaking if if we have traceability from the the you know the primary source right so here's the data use agreement that allows us to share this data here's how we've interpreted that that that line in the data or in the data use agreement and here's how we've implemented in policy and here's the actual data elements that are being shared usually if we get to that level of explicitness people will will understand but sometimes people just say sometimes they well i don't care what the data use agreement says i'm not doing it because of xyz um so there's always the human factor but but we've had pretty good success where we've gotten down to that level of transparency but there are always um things that you can't get over but anything about yeah yeah the only thing i would add in a slightly different approach uh or different perspective is that you know of course we've come up against the the compliance checkers or the the scanners or whatever it is but like part of it is i know what i'm going to come up against and if i've proactively assess myself under that compliance checklist that they're going to use and i can show them like i know what you're looking for i checked myself against it here's what i found and how i addressed it and start to build trust and rapport and show positive intent it doesn't happen immediately but i have seen over time that that compliance that eventually they start to kind of realize that you're fighting the same fight that you've got the same positive intentions and that you're taking security seriously and if you can show it to them in a way that sort of co-ops them into the solution that the next time they may be a little more accepting or willing to move things faster or they start to realize that this group really knows what they're doing let's give them a little more room to operate all right let's give a big thank you to our speakers
Trust, speed, clarity
Future-proof your data
Trust is fragile
Streaming data insights
Beyond technical solutions
Public sector challenges














