- Addressing 30 Petabytes of Data and 10,000 Applications
- Balancing Data Democratization with Enterprise-Grade Security
- Strategic Choices for BI, AI/ML, and Transactional Workloads
- Navigating Performance Trade-offs and Vendor-Specific Challenges
In an era where data-driven decision-making is paramount, organizations grapple with vast, disparate datasets and a myriad of applications. The World Bank, with its mission to end poverty and boost shared prosperity, faces this challenge on an unprecedented scale, managing over 30 petabytes of data across 10,000 applications. This session delves into their journey to establish a unified enterprise data serving layer, leveraging Databricks as the central nervous system for their data ecosystem.
The World Bank's cloud migration journey highlighted a critical need to eliminate data silos while fostering an open and democratized data framework. Their strategy focused on six key areas: removing silos, enabling application teams to use preferred tools, centralizing data with out-of-the-box capabilities like data protection and integration, and ensuring scalability and performance. After evaluating market options, Databricks emerged as the platform of choice due to its agility, security, integration capabilities, and rapid time-to-market for new solutions. It became the cornerstone for bringing data in, exposing it through Unity Catalog, and making it universally available.
Addressing diverse use cases, the World Bank identified three primary scenarios: data products, self-service analytics (using tools like PowerBI, Alteryx, and Tableau), and transactional applications requiring high-scale data consumption. Meeting these demands meant handling hundreds of simultaneous users, processing billions of records, and achieving report render times under 15 seconds. For applications, supporting JDBC, ODBC, and REST APIs was crucial, alongside enterprise integration for active directory, permission management, object-level security, and masking. These stringent technical requirements underscored the need for a robust, integrated platform that could deliver governance and security by design, rather than as an afterthought.
Ivana detailed the architectural implementation, showcasing an open data lake at the base, topped by a Delta Lake uniform layer for ACID transactions and time travel. Unity Catalog provides centralized governance, discovery, and access control. Above this, three serving tiers—quasi compute (interactive clusters, SQL warehouses), serverless compute (serverless SQL warehouses, notebooks, model serving), and API services (Delta Share, Azure API Management)—cater to different needs. Six specific serving engines, including SQL Warehouse, API Management, Delta Share, Federation, OLTP, and Model Serving, are aligned with distinct user personas and application types. Key lessons emerged around optimizing SQL warehouses (serverless vs. classic, scale-up vs. scale-out), navigating PowerBI driver limitations, and strategic caching to avoid performance bottlenecks.
For low-latency transactional applications, the current approach involves creating data copies in SQL Server with a Redis cache layer, acknowledging the trade-off for millisecond performance. Data virtualization through Databricks federation offers a unified view of disparate data sources, but comes with considerations like limited engine support, performance overhead from additional data hops, and increased cost. Ultimately, the World Bank's experience emphasizes that while a truly one-size-fits-all centralized data layer is challenging, covering 80-90% of scenarios with a robust platform like Databricks, and strategically tailoring solutions for critical applications, yields significant benefits in data discoverability, accessibility, and readiness for future AI/ML initiatives. Careful attention to vendor-specific drivers and continuous monitoring are paramount for success.
“If you have your data available and discoverable and accessible in a centralized, protected platform that is enabled for AI, ML, and all these futuristic and forward-looking applications that you want to build, then probably you can sacrifice here and there performance on an integration scenario.”
- Ivan Donev, Data Platform Architect, The World Bank




