Use case · Big data & analytics
Big data analytics
Big data analytics processes datasets too large for a single machine, using distributed frameworks across a cluster. It needs storage with capacity and throughput for large data, many cores and ample memory across nodes, and infrastructure sized to sustained processing. Dedicated bare metal is economical for steady big-data work, and EU hosting keeps large datasets under European sovereignty.
Key points
- Big data analytics handles datasets too large for a single machine, using distributed processing.
- It runs on clusters, with frameworks spreading work across many nodes.
- Storage needs both capacity for large data and throughput to feed the processing.
- Sustained big-data processing keeps clusters busy, which favours dedicated bare metal on cost.
- Large datasets of behavioural or user data make EU sovereignty a real consideration.
What does big data analytics need?
Big data analytics is the analysis of datasets so large that they exceed what a single machine can handle, which changes the infrastructure required from that of ordinary analysis. When data grows beyond the capacity or processing power of one server — into the volumes generated by large applications, extensive logging, sensor data, and similar sources — analysing it needs infrastructure that can store and process data at that scale. This means distributed processing across a cluster of machines, storage with the capacity and throughput for large data, and the compute to process it, working together to handle data too big for a single system.
The defining challenge is scale, and meeting it shapes every part of the infrastructure. The data must be stored across enough capacity and served at enough throughput; the processing must be distributed across enough compute to get through it in reasonable time; and the whole must be coordinated to work on one large problem. Big data analytics infrastructure is therefore built around distributing storage and processing across a cluster, sized to the volume of data and the analytical work. This page looks at big data analytics broadly; specific frameworks and systems for it, such as distributed processing engines and data warehouses, are covered on their own pages.
Distributed processing: analysing data beyond one machine
The core technique of big data analytics is distributed processing — spreading the work of analysing large data across many machines that process it in parallel. Because the data is too large for one machine to process in reasonable time, it is divided across a cluster, with each node working on part of it, so that the combined power of the cluster handles the whole. Frameworks built for this coordinate the distribution of data and computation across the nodes, letting large analytical jobs run across a cluster as if on one large system. This distributed processing is what makes analysing big data tractable.
Distributed processing frameworks — such as distributed processing engines and the broader ecosystem of big data tools — provide the means to express and run analytical work across a cluster, handling how the data and computation are spread and combined. These frameworks run on the cluster's infrastructure, coordinating its nodes, so the infrastructure beneath them must provide the compute, storage, and networking the distributed processing needs. Hosting big data analytics therefore means providing the cluster on which these frameworks run — enough nodes, with the compute, storage, and interconnect to process large data in parallel. The frameworks and the infrastructure together make distributed analysis of big data possible.
Storage for big data
Big data analytics needs storage that provides both capacity, to hold large datasets, and throughput, to feed the data to the processing fast enough. The datasets are large by definition, so ample storage capacity is required simply to hold them; and processing large data means reading large amounts of it, so the storage must deliver high throughput, often to many nodes at once, or the processing waits on data. Big data storage is therefore about serving large volumes at high aggregate throughput across a cluster, which is a different problem from single-server storage.
Meeting this often means distributed or high-throughput storage that serves the cluster's nodes with the data they process. Storage that spreads data across the cluster, or provides high aggregate throughput to it, keeps the processing fed; capacity accommodates the large datasets; and the storage is accessible to the nodes doing the work. Underprovisioning storage capacity or throughput bottlenecks big data analytics, leaving compute waiting on data, so storage is a first-class part of the infrastructure. We provide storage sized for the capacity and throughput big data analytics needs, so that the cluster's processing is fed by storage that can keep up with large-scale analytical work.
Compute: cores and memory across the cluster
Big data analytics needs substantial compute — many CPU cores across the cluster's nodes, and ample memory — because processing large data in parallel uses the combined power of many machines. The total compute of the cluster, aggregated across its nodes, determines how quickly big data can be processed, so clusters for big data analytics bring together many cores across many nodes. High-core-count processors, such as AMD EPYC, suit this by providing many cores per node, multiplied across the cluster. Some big data processing is also memory-intensive, holding data in memory for faster processing, so ample memory across the nodes matters.
Sizing the compute means providing enough nodes, with enough cores and memory each, to process the data at the scale and speed required. The right size depends on the volume of data and how quickly it must be analysed, and how much of the processing is memory-intensive. Some frameworks process data largely in memory for speed, benefiting from generous memory; others stream through data with less memory pressure. We size the cluster's compute — cores and memory across the nodes — to the analytical workload, so that the distributed processing has the aggregate power to handle the data at the scale big data analytics demands, without nodes being starved of the resources their part of the work needs.
Batch versus streaming analytics
Big data analytics comes in two broad modes — batch and streaming — which place somewhat different demands on infrastructure. Batch analytics processes large datasets in bulk, running jobs over accumulated data to produce results; it is about throughput, getting through large volumes, and can run when scheduled. Streaming analytics processes data continuously as it arrives, analysing it in near real time; it is about handling a constant flow and producing timely results. Many big data operations use both — batch for large periodic analyses, streaming for real-time needs.
The infrastructure implications differ between them. Batch processing needs the throughput to get through large volumes when jobs run, and its load may be periodic, concentrated when batch jobs execute. Streaming needs to handle data continuously as it flows in, with the capacity to keep up with the incoming rate at all times, and to process it with low enough latency for its real-time purpose. Sizing infrastructure for big data analytics means accounting for whether the work is batch, streaming, or both, and providing for the pattern each imposes. We size the cluster to the analytical modes it serves, so that batch throughput and streaming continuity are each supported as the workload requires.
Cost, utilisation, and bare metal for big data
Big data clusters represent substantial infrastructure, so their cost and utilisation matter, and sustained big data work tends to favour dedicated bare metal on economics. A cluster kept busy with ongoing analytical work — regular batch jobs, continuous streaming, steady querying — is well-utilised, and for such sustained use, dedicated hardware at a fixed cost is usually more economical than on-demand rates for capacity that is busy most of the time. Because big data clusters are large, the cost difference between dedicated and on-demand for sustained heavy use can be considerable, which is why steady big data work often runs on dedicated bare metal.
Bare metal also suits big data analytics for performance, since distributed processing over large data extracts a lot from the hardware, and dedicated hardware without virtualisation overhead delivers its full performance. On-demand capacity retains a role for genuinely intermittent big data work — occasional large analyses that do not justify a standing cluster — where paying only for the time used suits the sporadic pattern. The economical design depends on the utilisation: sustained big data work on dedicated bare metal, intermittent work on on-demand capacity. We size big data infrastructure to the pattern of the work, favouring dedicated bare metal where the utilisation makes it economical, and will say so honestly where on-demand capacity would suit intermittent work better.
Sovereignty for big data
Big data analytics often works with large datasets derived from real activity — user behaviour, events, transactions, and other data about people and their actions — which can be sensitive and subject to data-protection requirements, making sovereignty a real consideration. The very scale of big data means these datasets can contain a great deal of personal or sensitive information in aggregate, and where they are stored and processed governs a large volume of such data. For analytics on data about European users, or data subject to European protection, where the big data infrastructure runs is a data-protection question.
This is where EU-hosted infrastructure serves big data analytics on sensitive data. VV Internet Hosting is incorporated in the Netherlands, within the EU, so big data infrastructure hosted with us runs under European jurisdiction and outside the direct reach of the US CLOUD Act. For analytics on personal, behavioural, or regulated data — as much big data analytics is — keeping the large datasets and their processing in the EU keeps them under European law, rather than exposing a large volume of sensitive data to foreign legal access. For organisations to whom the sovereignty of the extensive data their analytics handles matters, keeping big data analytics in the EU addresses it at the scale big data involves.
Where VV Internet Hosting fits — and where it does not
We host dedicated bare-metal infrastructure for big data analytics: clusters of high-core-count nodes with ample memory, storage sized for the capacity and throughput large data needs, and the interconnect distributed processing requires — in EU datacenters under European jurisdiction. This suits organisations running sustained big data analytics that want performance, economical dedicated hardware, and European data sovereignty, on infrastructure they control. For that, we are a strong fit, and we will size the cluster's compute, memory, and storage to the analytical workload and its scale.
We are clear about our limits. Intermittent big data work — occasional large analyses that do not justify a standing cluster — may be better served by on-demand capacity, and we will say so. If you need a hyperscaler's fully managed big data platforms and their proprietary analytics services and ecosystem, that is a different kind of provider than we are. We provide the dedicated, sovereign infrastructure on which you run your big data analytics and its frameworks, on hardware you control — suited to organisations that want control, economical dedicated hardware for sustained work, and European sovereignty, and not a fully managed analytics platform. If our model fits your big data work, we can host it well; if not, we will point you toward what fits.
Questions
Big data analytics, answered plainly
Common questions about hosting for Big data analytics.
What makes big data analytics different from ordinary analysis?
Scale. Big data analytics handles datasets too large for a single machine to store or process in reasonable time, so it uses distributed processing across a cluster — dividing data and computation among many nodes that work in parallel. This need to distribute storage and processing across a cluster shapes the whole infrastructure, unlike single-machine analysis.
Is dedicated or on-demand infrastructure better for big data?
It depends on utilisation. Sustained big data work — regular batch jobs, continuous streaming, steady querying — keeps clusters busy, which favours dedicated bare metal at a fixed cost, more economical than on-demand rates for large clusters busy most of the time. Intermittent work — occasional large analyses — can suit on-demand capacity, where you pay only for the time used.
Why does sovereignty matter for big data analytics?
Big data often works with large datasets derived from real activity — user behaviour, events, transactions — which can contain a great deal of personal or sensitive data in aggregate. Where it's stored and processed governs a large volume of such data. Keeping big data analytics in the EU keeps it under European jurisdiction, outside the direct reach of the US CLOUD Act.
Planning Big data analytics infrastructure?
We host dedicated, EU-sovereign infrastructure sized to your workload — and we will tell you plainly when something else fits better. Tell us what you're building.