The Centers for Disease Control and Prevention (CDC) has unveiled a transformative accelerator methodology that is reshaping how big data is visualized and utilized in public health. Spearheaded by John Boer and his team, this initiative, known as CDC Data Hub, has overcome significant hurdles in data engagement, governance, and performance, leading to dramatic improvements in delivering critical health insights.
“You can have the best generative AI in the world if you don't store these things in a way that's meaningful and useful then you really lose a lot of its value.”
- John Bowyer, Data Architect, CDC
Struggling with slow dashboards or data chaos? The CDC shares its groundbreaking accelerator methodology, revealing 10 process and technology advancements that revolutionized their public health data visualization. Discover how they achieved 10x performance improvements and streamlined data products for critical insights.
and Welcome to our presentation Big Data visualization in public health just letting folks roll in here I'm John I'm John Boer I'm from The cdc's Office of Public Health data and surveillance and I'm from the federal government and I'm here to help so if you find those words a little scary I hope they'll be a little less scary by the end of this presentation um I've been working seriously on a lot of these innovations that you'll see today very hard for the last five years with or over five years with over a hundred people across six centers and offices at the CDC and over that time We've ran into a lot of challenges and learned a lot of lessons and I'm happy to share those with you guys today this is the first time I presented these findings publicly and so um I'm excited to do so and when on the team on those on those lessons we're going to focus on two main areas Big Data visualization in public health and building faster and more reliable data
products so a little disclaimer as we begin the presentation it's important to note the context and the scope of the information that I'm presenting today the findings and the conclusions presented are those of myself and the contributing authors and they don't necessarily reflect the official position of the Centers for Disease Control and prevention the data and examples are based on specific CDC use cases and don't necessarily con constitute an endorsement of any product so what have I been doing and my
our team been doing the name of our project is CDC data Hub and CDC data Hub is a One-Stop shop for data at the CDC it has over 50 unique data assets from 13 different providers and it covers visits Labs medications costs Etc and it's used by 450 users across 18 CDC offices and centers and what do they do with that they produce a lot of research output to help Public Health we facilitated over 900 Publications since 2009 contributing to 46,000 citations in scientific literature and we've done all of of this the whole Five Years on the data bricks platform and it's been an awesome journey and I can't think of a place I'd rather be than where I am at the CDC and it's it's really fulfilling to do this work so presentation objective and
agenda uh so we're going to explore an accelerator methodology that we've just implemented this year in 2024 on top of the the base CDC data Hub architecture and that accelerator methodology solves 10 specific challenges that we faced and we'll step through those 10 challenges today and highlight how we addressed them and we'll break we break those 10 challenges down into two parts first we'll start with people and then we'll go into technology and why should it matter to you guys while a lot of this has to do with public health I'm going to try to focus on the areas that apply to you the broader non Public Health audience data products Big Data visualization in general and use our implementation of an RSV dashboard as a specific example and those are my bullets so a little quick bit about me my name's John and I've been working on data for over 30 years I started out in the 80s with Punch Cards in Fortran at UG and then moved on in the later 90s to work for Fortune 500 companies modernizing main frames and using SQL for when it was new and new in the SQL Server sense and then moved on to the dot boom and bust in the early 2000s started working with Pharmaceuticals in the in the two early 2000s as well spent about 10 years at Microsoft before I spent the last five years at the CDC is both a contractor and a a full-time employee FTE so I'm went to Georgia and I live in Tuscaloosa Alabama and if you know anything about the SEC it's uh it's a pretty hard place to find happiness for a Georgia grad but I'm super I couldn't again I couldn't be there's nowhere else I'd rather be than where I'm at and a little more about CDC data Hub which is the project that I'm working on um again today we're going to cover the five improvements on the left for the process people side of things and then move on to the five technology s u advancements that you see on the right and I think we should just sort of dive right into it and and get started on the people side of things
so on the people's side process modernization so process Improvement number one it was that the hardest part of sort of helping people with visualizations was having them find the visualizations and so this was the um the the number one challenge was in what I'm calling the the engagement Challenge and you can kind of think of it as sort of dashboard Drifters if you were to be able to see this presentation and look closely at the the right hand image there you would see that there were just a list of reports in standard powerbi format some of them said old some of them said broken and some were the production reports that our users were supposed to find and what what we needed was a site that was integrated with our internet and was uh searchable discoverable and it also gave confidence that there wasn't intermingled um in a in development work separated from a production work another lesson learned from the covid response was that this architecture that you're seeing while we just implemented in CDC data Hub it's been around for five years so what was it doing it was actually used in the co response to send uh reports to the White House and other areas and those external V external customers of the CDC see don't necessarily log in and look at a powerbi dashboard so being able to produce high quality Rich PDFs Excel on a daily basis that were in an automated way go to their emails and to their SharePoint sites with the you know with an Excel file named with the date of the current day or week were very important during the response and so we've also included those methods in this in the in this so this is what we've got now we've got a a nice new internet site and it's got our reports embedded in it and they have not just powerbi dashboards but they also have powerbi report Builder exports in Rich formats and so that's sort of um opportunity number one for for optimization in the process side of things so the second process Improvement
number two is agile user stories and recipes and so I mentioned that there were hard lessons learned and this is probably the hardest personal Lesson Learned was that I worked really hard over those first three or four years and created a 500 page sop document on everything we should all do for to to solve these problems and challenges and for anyone embarking on such an Endeavor I would just caution to not be overly optimistic about people's response to a 500 page document of any sort um it doesn't really matter the quality because they're probably not going to open it and it's kind of it was it was a little bit humbling um to work that hard efforts not always related to success sometimes um resilience and adaptability are more important so the lesson that I learned from that was that we needed to change and so we created these things called recipes which are just really specific targeted um Step by-step guides and we got these idea the idea of a recipe from the vendors vendors uh data bricks has a really good concept they use a machine learning but then Eli and posit also use this concept of recipes to and and they and they work and so that's that's really rewarding because some of the same folks that I've been sending this big document to for years now I just they need to on board a user to GitHub or they need to cicd um recipe and they're liking it they're it's it's almost shocking sometimes I'm like wow success but it's not necessarily like a straight line Journey you'll see in these things um to get there U um so that's that was process Improvement number two so process number impr Improvement
number three was Mach the machine readable requirements and this is a a real challenge because the way that people typically do these are sticky notes and then if they're adapted SQL they'll just put the requirements right in the SQL that they write so there's this sort of uh vague Vortex where the requirements that that we need to be in a document that's um compliant don't exist in the format that a machine can understand and if they're on a sticky Noe a lot of times what they'll say is that I need a longitudinal analysis of any data source over any time for any reason by next week and that's a good ambitious goal and it is good to get users desires on paper but what we really challenged with was creating a format that a machine can understand and so what we've done is we've created an Excel template the latest version of it and it captures requirements in a gridlike form that are machine readable and human readable and this is and this there obviously requires a little bit of of training of our users which we have to engage in but training users on Excel to capture their ideas is still a much more efficient and and beneficial process than training them on SQL and having it hard hardcoded so that's um the that Excel though that is entered is uploaded into the data Lake and it does end up in data bricks and is drives are is literally not only is it machine readable but we do actually read it and drive our processes off of the machine readable requirements so you're in a big you're in a big data visualization talk and so I imagine some of you guys want to see some big data visualization so woohoo so so thank you guys so um so process Improvement number
four streamlined content dashboard interface so what you're seeing is again this concept of templatized reusable uh work the template that you're seeing is a starting point they don't have to use this but for each project we follow a certain data do data structure which I'll go over to later and at the top of that data structure is an app in this case you're looking at a explor dashboard and then beneath that there's a second level subcategory you might wonder where's the category this is a powerbi report that you're looking at and the tabs down at the bottom which you can't see are the categories so once the user can drill down through the the app level which is data driven to the subcategory which is data driven they can get to what we call a view byby which is their Dimension level partitions for Big Data visualization and this allows them to to view anything by age I think right now we on this report we have somewhere between 12 and 20 dimensions and then for each of those Dimensions you can drill down further and choose one or more of the um subcategories and then you can also drill them all by time and then the metrics are across the top so the point of this visualization is not so much just the ux which is is is admittedly pretty descriptive statistics generically based but to to Really set the basis for data model driven a data driven data driven approach to uh to visualization to because at the end of this you will see that there's 10 times Improvement by the optimizations that we made we had an initial version of this dashboard and then we have the current version and the current version does a lot more with a lot less resources and a lot better user
experience so this is the process Improvement number five is I think a generic enough process Improvement that we've probably all experienced it inside or outside of technology which is that it's what I'm calling the the um the phantom facts challenge where project reports or project plans those Gant charts that leadership may see may not necessarily reflect exactly the details that you see in your gr dashboard they're probably not necessarily created a and automated way from jir and if they do match and even if they're even more correct than jir there's not a way to drill down into your details so there's this sort of Office document PowerPoint leadership world with things and then there's the J tactical real world on the ground point of things so what we've done by that to to help address this process challenge is integrate office in jir with an API and and make it so that we have PowerPoints that allow the use the in Excel documents that allow our leadership to drill down and see the results and also if you're in a standup it's sometimes a lot more efficient to just you know get the data quickly entered in filter in Excel and then enter in a new jot task at the end of the standup so these improvements have made these are that's these five process improvements have made substantial probably even more impact on our technology stack than our technology stack so on to the technology stack but I'm biased this is still my favorite stuff so um um we will get into the um the technology advancement number one now which is that this is our common data model so if
you think of this as a a pipeline factory and we're building pipelines this is this is the this is sort of the schematic of how we we build our pipeline structures and what we had before was we um we had the reason that report rep was so slow um we it was kind of a molasses report before and we're 10 times faster now that visualization that I showed you guys earlier was because the initial report that I that that I began to use had 400 columns and to add another metric they would they would add columns for the combination of the metrics and the dimensions had six rows and 400 columns so I call that the the you know the pancake stack dmma here where again it was almost like my my document you sometimes you just got to revisit things and say this is not the right approach and you need a normalize data model and in some ways again it's a blessing because it forces you to re to rethink everything and that's what you'll see here is our data model for visualizations so you'll see that if you look at this you've got the app on the right and the view by categories and it's all data driven in Normal but in the middle of this is we've used a lot of common there's so many there's so many common data models in the world and so you might think we have the the sort of rotten ta Tomatoes of of aggregator of of data models we use omop we use fips we use um a lot of different data models and we we sort of standardize those and carbonize those on the data bricks platform and it's important to note that you can also put different data in in the same data model so while we may use an omop schema and I'm throwing out a lot of acronyms that I'm not going to necessarily go through each one of them you can just think of these as database standards and but we we've sort of harmonized them together I guess that's the key Point into these structures that you're seeing on the screen and the the bottom structure here is for a pipeline Factory this is sort of the basement of our of our pipeline Factory this is the these are the data structures that we use to drive the ETL processes and interesting enough those same Excel spreadsheets that we talk about machine readable requirements are used in each of these areas as well so if a user goes to configure a new job or a new workflow or a new bronze data set uh and then this is those spreadsheets and machine readable requirements funnel into these data structures so that's technical technology advancement number one which is our
pipelines so on to technology advancement number two which is our data conversion so I'm calling this technology challenge um conversion chaos which is and I've got a little we've seen this talked about a little bit during the even during the keynote this morning but I actually put some visuals about how hard it is when you go to different like I said we have all those different vendors supplying the same EHR electronic health record type data and they're all in different formats not just physical file formats but the structures are different and that's super hard for our scientists and it's super hard for the people consuming our work if we replicate the problem and give all those outputs out so while data bricks has done a masterful job of harmonizing the the um the data sharing process what we're talking about here as a universal data connector is having a universal data structure and a some standard formats so that the the so that the the people consuming our data get their pivots with you know geography means the same thing and in a standard way and that they get the data and pivot format if if that's the most uh appropriate way for them to receive the data another way that we standard iing this data conversion is I think we underutilized and and and I would really want to brag on the data breaks platform for this if you guys haven't been using datab break SQL or if you are using datab break SQL I super encourage you to look at parameters and the keynote this morning she just did a couple brackets and she just entered text but if you if you're using datab brick SQL enter a couple brackets and choose the drop down option and see how awesomely powerful it is to have users use parameterized SQL so what we can do with this is we can have one generic SQL statement create 50 to 60 different views in an automated way just accepting a different parameter and the output of those 50 or 60 different views if we have a you know a standard long skinny code lookup table with state name you know just code decode when we go and pivot it we can actually make the state star table have the name state code in state based on the parameter that comes in and if it's you know gender it says gender name and gender because it's using those parameters to create the column names and it's a tremendously powerful way to reduce complexity and still give people The Best of Both Worlds in a data model and I haven't seen enough of those features um evangelized here because they've done a great job on them and again this is is the this is the the the common data model structure that I just talked about for codes and decodes it's based on omop but it's all of you probably have some sort of value set or look up whatever your name is for the same
thing so on to technology advancement number three data governance expectation tracking and tools and if you felt a little oomph in my voice here I guess I'm a little passionate about this one um um I'm calling um I'm calling this the opaque Oracle problem but there's a key thing first you need to have machine readable requirements they can drive your requirements but you need to have expectations that you Monitor and those those expectations need to be captured so a manager may say they want to have a longitudinal analysis that does anything any time but those aren't measurable so we have to have the same kind of machine readable rules that we did for the requirements did for our testing and that's part of what this slide is about but the part that I'm most passionate about is the synthetic data part of it because I think the biggest piece of ethical AI is that if somebody tells you a secret you shouldn't share it and the best way of not sharing Secrets is synthetic data in my mind that doesn't mean that we never test with real data but we should be able to test with synthetic data and if there's an I'm I'm sure no one in this room has ever heard oh you can't see my data because it's private or secure but I've heard that before and if we all use synthetic data or there was some sort of mandate for that sort of thing then init in day one we can share our data both internally and externally and what's a now a a private problem for me at the CDC can be a public problem for you guys as Citizens because the data is available on kaggle for you to start helping solve public health problems with synthetic data so maybe I'm putting too many too high expectations on you as an audience but those are the kind of dreams I have for us is that you know we'll start being able to capture these requirements better and then share them with you and then help engage the citizens to to help us solve these public health
challenges so there's been a lot of talk about data about data products and everybody has their own definition and this is my definition since I'm on stage so um I'm for me the challenge that the data product is is solving is the metadata mystery and so a lot of times if you think of the pipelines that we talked about as sort of the the structure and the data as the gas then the data product is sort of the packaging this is what the customer sees and this is how they engage with it it's like they said in the keynote or somewhere that the it's a product life cycle where you know you have a start date and an end date of your product and you have some governance rules around it and you need to catalog solves some of this but it doesn't prescriptively give you rules and what you'll see consistently throughout this expectation is or this presentation is that we're trying to sort of raise the bar in terms of structure and guidance that we provide folks so for us a data product means that it's has a it's a database or a data source and then those data sources can have multiple um data sets or tables in a schema and then they can map to an individual data set and their package in a way that is um is Meaningful to our users and they can um so an example to make this a little more real world is that these analytical pivots that we're talking about are data products they're part of the data product and so for US Gold doesn't mean putting it in a flat table and having the user create their own metrics and Metric definitions we're actually defining what the metric definitions mean what the dimension definitions mean in our data
products so we're about done so on the on the the five of five standardized data um data product life cycle uh technology advancement number five and this one I'm calling the reinvent rocket and so a lot of times what you'll find on these projects is that we were where you recreate things from scratch or at least that's what we did and you don't necessarily have a consistent ETL process and so just like we talked about the the pipelines for the structure and the gas for the data but you can think about this this standardized product life cycle as the pump the pump the gas through the pipelines in our Factory framework and the key components are down here at the bottom we call it the ideas uh framework and so that means it starts with Ingress or ingestion the data comes in from the internet and it goes to a data stores on on a raw data storage and then it ends up on the far right with the S on the serving and storytelling and the the parts in the middle will go through so it's it comes into Ingress it goes into our data Lake storage when we do a data load and it goes into our bronze table and these are actual function names in in our in the code there's 50,000 lines of code to to drive all of this magic we've been talking about and then it goes into this process process data load section then it goes into the enrichment stage so once it's into a raw bronze table we can add geocoding we can extend it and extend the data with other additional elements then we can do our analysis phase which is all the fancy pivoting and things that we talked about earlier um and then we can do the Final Phase which is our s our our storytelling phase so that's the end of our top 10 list so
um you again we've made significant strides increasing the CDC data Hub and I'm super excited about what we've been able to achieve and I'm super excited to have you guys here and uh in in with me so on the data process improvements the uh this these slide this slide goes [Music] [Music] over the people side of things that we covered and while I don't have as much quantitative measures of the I can the 10 told Improvement is the the technical provable Pro performance improvements but I have had epidemiologists come up to me and say this this new process is a thousand times faster than the old process and when I mathematically think about it takes about 3 minutes to do the spreadsheet now and it took about three days to do it before I don't know that she's totally wrong so a lot of again the reason that we may have spent more time in a technology talk on process improvements is because you might get more Pang from your butt from those but so what are the types of techn things that we went over we went over how can we we reduce hard coding in our code we've gotten rid of the global reference we've introduced Global reference data standards and we've reduced copy pasting in SQL I also want to point out on the powerbi side we've made dramatic reductions so if you're putting 50 buttons on a powerbi report and you can put five tabs not only is it that much faster for the developer it's actually responds that much faster so we've made dramatic simplifications in the techn ology stack as well in powerbi and also the amount of calculations that are done in the front end the the the lines of Dax code has just been has been totally dwarfed by data briak routines now and the net result is that we've been able to make our technology perform impr provably 10 times faster and the reports this is made before we were at an impass you couldn't scale not only does it perform 10 times these these differences in performance would only grow as the number of columns in obstacles gr so we've got a platform for the future now and the these these um it's minor connector went minor connector went away all right scary so thank you guys so again we've reduced the file size we've reduced the the application complexity and we've also made it you know the net result is that the thing is faster to load but it also is a huge visual user interface experience because when it's that much if you guys have seen a chart crawl versus a chart that's that's flows it's a much better user experience so what did we over today we went over dashboard Drifters not being able to find dashboards we went over my personal death by docs experience where the documentation is too big we went over the vague Vortex experience Pro challenge we went over the Molasses matric when the reports too slow we went over the Phan of facts problem where project managers have reports that aren't substantiated necessarily by the facts we would like to see we have the pancake stack problem with a the the the data table's too wide we have the conversion chaos problem where there's not consistent Universal adapters to convert data we went over the opaque Oracle where we don't have defined quality expectations and we don't necessarily have the synthetic data or if we have it we're not sharing it with everybody like you and we have the metadata mystery problem which is where we don't catalog our analytical products with at the metric and dimension level and we have the reinvention rocket where we do the same idea steps over and over like the ideas new and it's it's really a lot of the same things that can be used so that pretty much wraps it up guys so that those are our prod those the key takeaways are that there's the process modernization and the technology modernization and the RSV dashboard I do want to say that this is there's so much more what you've seen is a ly descriptive statistics and we can do predictive statistics and we can do prescriptive statistics but you can have the best generative AI in the world if you don't store these things in a way that's meaningful and useful then you really lose a lot of its value because your leadership doesn't need an interactive session like with the chat GPT when when it comes to reporting their the the the overall annual report you know you need that needs another level of diligence and that's the kind of diligence I've tried to show you guys today so that in that concludes our presentation and I have some other slides here or we can do questions if you guys have some questions I'm not sure how we're doing time
Recipes for success!
Data model driven UX!
Documentation overload is real!














