Use case · Big data & analytics
Data warehouse hosting
A data warehouse is a central repository of integrated, structured data from across an organisation, optimised for analytical queries and business intelligence. It needs data integration, storage for historical data, and compute for fast analytical queries. Self-hosting an open warehouse engine on dedicated EU hardware gives control, performance, and European sovereignty over the organisation's integrated data.
Key points
- A data warehouse is a central repository of integrated data optimised for analytics and BI.
- It brings data together from many sources through integration processes (ETL or ELT).
- Its data is structured and modelled for analytical queries and reporting.
- It needs storage for historical data and compute for fast analytical queries.
- Self-hosting an open engine on EU hardware keeps the organisation's integrated data sovereign.
What is a data warehouse, and what does it need?
A data warehouse is a central repository that brings together integrated, structured data from across an organisation, organised and optimised for analytical queries and business intelligence. Rather than data being scattered across the many systems that produce it, a data warehouse consolidates it into one place, structured for analysis, so that reporting, dashboards, and analytical queries can draw on a coherent, integrated view of the organisation's data. The warehouse is where an organisation's data is gathered and made ready for the analysis that informs decisions, which is a distinct role from the operational databases that run its applications.
Providing a data warehouse means supporting this consolidation and the analytical querying it exists for. The warehouse must integrate data from many sources, store it — often including substantial history — in a structured form, and serve analytical queries over it quickly enough for reporting and business intelligence. This calls for infrastructure with storage capacity for the accumulated data and compute for analytical queries, together with the processes that load and organise the data. This page looks at data warehousing broadly; specific analytical engines that can power a warehouse, and the broader practice of big data analytics, are covered on their own pages.
Integrating data: bringing it together
A defining function of a data warehouse is integrating data from across an organisation's many systems into one coherent repository, which is done through processes that extract data from sources, transform it into a consistent form, and load it into the warehouse. These integration processes — whether transforming data before loading it or loading it and transforming it within the warehouse — bring together data that originates in different systems, in different forms, into the structured, consistent form the warehouse holds. This integration is what makes the warehouse a single, coherent view rather than a copy of scattered, inconsistent data.
The integration is central because the value of a warehouse lies in having the organisation's data together and consistent, ready for analysis. Data from different systems must be reconciled — made consistent in structure and meaning — so that analysis across it is valid, which the integration processes accomplish. These processes run regularly to keep the warehouse current as the source systems produce new data. Hosting a data warehouse therefore involves supporting these integration processes, which run on infrastructure and load the warehouse, as well as the warehouse itself. The infrastructure must accommodate both the integration work and the storage and querying of the integrated data.
Structured, modelled data for analytics
Data in a warehouse is structured and modelled specifically for analytical querying, which distinguishes it from the way operational databases structure data for transactions. Warehouses commonly organise data using models designed for analysis — arranging it so that analytical queries, which aggregate and summarise across large amounts of data, run efficiently and are natural to express. This modelling shapes the data around the questions analysis asks, rather than around the operations an application performs, which is why a warehouse's structure differs from that of a transactional database.
This analytical modelling is part of what makes a warehouse effective for reporting and business intelligence. By organising integrated data in a form suited to analysis, the warehouse lets analytical queries run well and lets the data be understood and explored for insight. The modelling is a design activity, done as the warehouse is built, that reflects how the data will be analysed. Hosting a warehouse supports data organised this way, on infrastructure suited to serving analytical queries over it. Understanding that a warehouse's data is modelled for analysis clarifies why it is a distinct system from operational databases — built, structured, and hosted for analytical querying rather than transactions.
Query performance for analytics and BI
A data warehouse must serve analytical queries quickly, because it exists to power reporting, dashboards, and business intelligence that people use to understand the organisation and make decisions. Analytical queries scan and aggregate large amounts of the warehouse's data, and their performance determines how responsive reporting and dashboards are and how freely analysts can explore the data. Slow analytical queries make a warehouse frustrating to use and limit the analysis done on it, so query performance is central to a warehouse serving its purpose.
Achieving good analytical query performance depends on both the warehouse engine and the infrastructure beneath it. Analytical engines — including column-oriented systems built for exactly this kind of querying — process analytical queries efficiently, and they run on infrastructure whose compute, memory, and storage speed determine how fast the queries complete. Providing a performant warehouse therefore means an appropriate engine on infrastructure sized for analytical querying — compute for processing queries, memory, and fast storage to feed the large scans analysis involves. We provide infrastructure suited to serving analytical queries quickly, so that a warehouse hosted with us powers responsive reporting and business intelligence rather than being held back by slow queries.
Storage and compute for a warehouse
A data warehouse needs storage with capacity for its accumulated, often historical, data, and compute to run analytical queries over it. Warehouses commonly hold substantial history — data accumulated over time to allow analysis of trends and changes — so storage capacity is a real requirement, alongside the throughput to feed analytical scans. Compute is needed to process the analytical queries, which can be demanding, scanning and aggregating large amounts of data. The balance of storage and compute depends on how much data the warehouse holds and how heavy the analytical querying is.
Sizing a warehouse's infrastructure means providing storage for the data volume, including its history, and compute matched to the analytical query load. A warehouse accumulating years of data needs the capacity to hold it; one serving heavy analytical querying needs the compute to process the queries responsively. Fast storage feeds the large scans analytical queries perform, and ample compute processes them. We size storage and compute to the warehouse's data volume and query load, so that it can hold the organisation's integrated data, including its history, and serve analytical queries over it with the performance reporting and business intelligence need.
Self-hosted versus managed cloud warehouses
A significant choice is between self-hosting a data warehouse using an open engine on infrastructure you control and using a managed cloud data warehouse run by a provider. Managed cloud warehouses are convenient — the provider runs the warehouse as a service — which is valuable, but they carry the familiar trade-offs: the organisation's integrated data sits on the provider's platform, often under a foreign jurisdiction; costs can be substantial and depend on the provider's pricing; and the organisation is shaped by the provider's ecosystem. Self-hosting a warehouse using an open analytical engine on infrastructure you control keeps the integrated data, the engine, and the terms in the organisation's hands.
The honest basis for choosing is what the organisation values. If minimising operational effort and adopting a provider's managed warehouse matter most, and placing the organisation's integrated data on that provider is acceptable, a managed cloud warehouse may suit. If control over the warehouse, predictable cost, and — importantly — keeping the organisation's integrated data on infrastructure it controls and under a chosen jurisdiction matter more, self-hosting an open engine delivers them, which is what we support. Organisations often choose to self-host their warehouse when the integrated data it holds — a comprehensive view of the organisation — becomes too valuable or sensitive to place on a foreign managed platform. That is the case we serve, while acknowledging that a managed warehouse's convenience is the right trade for some.
Sovereignty for warehouse data
A data warehouse holds a comprehensive, integrated view of an organisation's data, which makes it among the most sensitive collections of data the organisation has, and its sovereignty correspondingly important. Because the warehouse consolidates data from across the organisation, it concentrates in one place a great deal of what the organisation knows — including personal, business, and regulated data — so where the warehouse resides governs an especially significant collection of data. For an organisation subject to European data protection, or wishing to keep its integrated data sovereign, where the warehouse runs is a substantive concern.
This is where self-hosting on EU infrastructure serves sovereignty over the warehouse's data. VV Internet Hosting is incorporated in the Netherlands, within the EU, so a data warehouse self-hosted with us keeps its integrated data under European jurisdiction and outside the direct reach of the US CLOUD Act, on infrastructure the organisation controls. For a warehouse consolidating sensitive or regulated data — as most do, given what they integrate — keeping it in the EU keeps that comprehensive collection under European law, rather than placing it on a foreign managed platform. For organisations to whom the sovereignty of their integrated data matters, self-hosting the warehouse on EU infrastructure addresses it for the concentrated, sensitive data a warehouse represents.
Where VV Internet Hosting fits — and where it does not
We host dedicated infrastructure for self-hosting a data warehouse: storage with capacity for integrated and historical data, compute for analytical queries, and fast storage to feed analytical scans — running an open analytical engine of your choice, in EU datacenters under European jurisdiction. This suits organisations self-hosting a warehouse that want performance, control, and European sovereignty over their integrated data, on a foundation they command. For that, we are a strong fit, and we will size the storage and compute to the warehouse's data volume and analytical query load.
We are clear about our limits. We provide the dedicated infrastructure you run a self-hosted warehouse on; we are not a managed cloud data warehouse service that operates the warehouse for you as a proprietary offering. If you want a fully managed cloud data warehouse where the provider runs everything, that is a different kind of offering. We are the dedicated, sovereign infrastructure on which you self-host a data warehouse using an open engine — suited to organisations that want control, performance, and European sovereignty over their integrated data, and not a managed warehouse service. If self-hosting a warehouse on infrastructure you control fits your needs, we can host it well; if you want a managed cloud warehouse, we will be clear that is a different offering.
Related use cases
Questions
Data warehouse, answered plainly
Common questions about hosting for Data warehouse.
How is a data warehouse different from an operational database?
A data warehouse is a central repository of integrated data from across an organisation, structured and modelled for analytical queries and business intelligence — consolidating data for analysis. An operational database runs an application, structured for transactions. The warehouse is built, modelled, and hosted for analytical querying over integrated, often historical, data, a distinct role from operational databases.
What is ETL or ELT in a data warehouse?
They're the integration processes that bring data into the warehouse — extracting data from source systems, transforming it into a consistent form, and loading it into the warehouse (transforming before loading, or loading then transforming within the warehouse). They reconcile data from different systems into the structured, consistent form the warehouse holds, run regularly to keep it current.
Should I self-host a data warehouse or use a managed cloud warehouse?
It depends on what you value. A managed cloud warehouse is convenient but places your integrated data on the provider's platform, often under a foreign jurisdiction. Self-hosting an open engine keeps the data, engine, and terms in your hands, at the cost of more operational work. Organisations self-host when their integrated data is too valuable or sensitive for a foreign managed platform.
Planning Data warehouse infrastructure?
We host dedicated, EU-sovereign infrastructure sized to your workload — and we will tell you plainly when something else fits better. Tell us what you're building.