Use case · AI & GPU compute

ML platforms

An ML platform is the shared, self-hosted infrastructure an organisation's teams run their machine-learning work on — a common pool of GPU compute, environments, storage, and orchestration serving many projects and pipelines. Sharing resources improves utilisation and cost, and self-hosting in the EU keeps the whole organisation's ML data under European jurisdiction.

Key points

  • An ML platform is the shared foundation an organisation's teams run their ML work on.
  • It pools GPU compute across many teams, projects, and pipelines rather than siloing it.
  • Sharing a common pool improves utilisation, which lowers the cost of expensive GPUs.
  • It provides common environments, storage, and orchestration as a shared capability.
  • Self-hosting the platform in the EU keeps the whole organisation's ML data sovereign.

What is an ML platform?

An ML platform is the shared infrastructure and capability on which an organisation's machine-learning work runs — the common foundation that its teams, projects, and pipelines draw on, rather than each building its own from scratch. Where a single pipeline is one workflow, an ML platform serves many: it provides a pool of compute, a set of environments, shared storage, and the means to run pipelines and serve models, available to the whole organisation's ML efforts. The platform is thus the layer beneath the individual pipelines and projects, giving them a common, managed foundation to run on.

The value of thinking in terms of a platform is that it consolidates what many ML efforts need in common. Rather than each team acquiring and managing its own GPUs, environments, and storage — duplicating effort and fragmenting resources — a platform provides these centrally, so teams draw on a shared foundation. This consolidation is what a platform offers: a single, well-run base of ML infrastructure that many can use, which is more efficient and coherent than scattered, per-team setups. Hosting an ML platform means providing this shared foundation, sized and organised to serve an organisation's ML work as a whole.

Shared compute: pooling GPUs across teams and jobs

A central function of an ML platform is pooling GPU compute so that many teams and jobs share a common set of GPUs rather than each holding its own. GPUs are expensive, and if each team has dedicated GPUs, those GPUs sit idle whenever that team is not using them, while another team may be short of capacity — a wasteful arrangement. A shared pool lets the organisation's GPUs be allocated to whichever work needs them at a given time, so that the GPUs are used by whoever has demand, keeping them busier and spreading their capacity across the organisation's needs.

Making this work requires scheduling — allocating the shared GPUs among competing jobs and teams. The platform must decide which work gets which GPUs when, so that the pool is shared fairly and used well, without one team's work starving another's or GPUs sitting idle while jobs wait. This scheduling is part of what an ML platform provides, turning a set of GPUs into a shared resource that the organisation's work draws on as needed. The result is that expensive GPUs serve the whole organisation efficiently, which is a large part of why consolidating compute into a platform pays off compared to fragmented, per-team GPUs.

Common environments and tooling

An ML platform provides common environments and tooling, so that teams work on a consistent, ready foundation rather than each assembling their own. As discussed for deep learning, ML work depends on a coherent software stack of frameworks, libraries, and drivers, and getting this right is real work; a platform provides such environments centrally, so teams can run their work on ready, consistent stacks rather than each building and maintaining their own. This consistency also aids reproducibility, since work runs on known, shared environments across the organisation.

Providing environments and tooling as a platform capability reduces duplicated effort and fragmentation. Instead of every team solving the same environment problems independently, the platform solves them once and offers the result to all, so teams spend their effort on their models rather than on assembling infrastructure. Common tooling for running pipelines, tracking experiments, and serving models similarly benefits from being provided centrally. When we host an ML platform, providing for these common environments — the consistent, working stacks the organisation's ML work runs on — is part of the foundation, so that teams find a ready base rather than having to build one each.

Shared storage for datasets, models, and artefacts

An ML platform provides shared storage for the datasets, models, and artefacts that the organisation's ML work uses and produces, so these are held and accessible in common rather than scattered. Datasets are often shared across projects, trained models and their versions need to be stored and retrieved, and the artefacts of ML work accumulate; a platform gives these a common home, accessible to the teams and pipelines that need them. This shared storage means datasets do not have to be duplicated per team, models are kept where they can be found and served, and the organisation's ML data assets are managed coherently.

This shared storage has to serve the platform's varied needs: fast where it feeds training, capacious for the datasets and models held, and accessible across the compute the platform provides. Providing it as a shared resource, rather than per-team storage, keeps the organisation's ML data organised and avoids the fragmentation and duplication of scattered storage. When we host an ML platform, the storage is designed as a shared part of the foundation, sized for the datasets and artefacts the organisation works with and connected to the shared compute, so that data and compute together form a coherent platform rather than disconnected pieces.

Sharing within the organisation: scheduling and isolation

An ML platform is shared among an organisation's teams, which makes it internally multi-tenant — many users and teams drawing on common resources — even though the platform as a whole is dedicated to the one organisation. This internal sharing raises questions of fair scheduling and appropriate isolation: the shared GPUs and resources must be allocated among teams so that the sharing is fair and the platform serves everyone, and teams' work should be isolated enough that they do not interfere with one another, while still sharing the underlying resources. Managing this internal multi-tenancy well is part of running a platform.

At the same time, the platform being dedicated to the one organisation, rather than shared with outside parties, gives it the isolation of single-tenant infrastructure at the organisational level. The organisation's ML work runs on its own dedicated platform, not commingled with other companies' work, which matters for the security and sovereignty of its data, while within the organisation the resources are shared among its teams. This combination — dedicated to the organisation, shared among its teams — is characteristic of a self-hosted ML platform, and getting the internal sharing right, with fair scheduling and suitable isolation between teams, is part of what makes such a platform serve the whole organisation well.

Self-hosted platform versus managed ML platforms

A significant choice is between building a self-hosted ML platform on infrastructure you control and using a managed ML platform run by a provider. Managed platforms offer convenience — the provider runs the platform, and teams use it without operating the underlying infrastructure — which is valuable, but carries the familiar trade-offs: your data and work sit on the provider's platform, you depend on their pricing, availability, and terms, and you are shaped by their ecosystem. A self-hosted platform, on infrastructure you control, requires operating it but keeps the compute, data, environments, and terms in the organisation's hands.

The honest basis for choosing is what the organisation values. If minimising operational effort and adopting a provider's managed tooling matter most, and placing the organisation's ML work on that provider is acceptable, a managed platform may suit. If control over the platform, predictable cost at scale, and keeping the organisation's ML data on infrastructure it controls matter more — especially where that data is sensitive or subject to European protection — a self-hosted platform delivers them, which is what we host. Organisations often choose to self-host their ML platform precisely when the data across their ML work becomes sensitive enough that a provider-run platform is no longer acceptable. That is the case we serve, while acknowledging that for some, a managed platform's convenience is the right trade.

Cost efficiency through a shared pool

One of the strongest arguments for an ML platform is cost efficiency through sharing, because pooling expensive GPUs raises their utilisation. When GPUs are dedicated per team, each team's GPUs are idle whenever that team is not using them, so across the organisation much expensive capacity sits unused at any moment. A shared pool lets the organisation's total GPU capacity be smaller than the sum of what separate per-team allocations would require, because the shared GPUs are used by whoever has demand, smoothing out the idleness. Higher utilisation of expensive hardware directly lowers the cost of the organisation's ML compute.

For a platform serving sustained organisational ML work, this favours dedicated infrastructure kept well-utilised through sharing. A shared pool of dedicated GPUs, busy across the organisation's combined demand, is economical in a way that neither fragmented per-team GPUs nor on-demand rates for continuous use match, since the sharing keeps the fixed-cost hardware working. On-demand capacity retains a role for peaks beyond the pool's size. The economical design is usually a well-sized shared pool of dedicated hardware for the organisation's steady combined load, and we will size that to the real aggregate demand rather than to the sum of separate per-team peaks, which is where consolidating into a platform saves the most.

Growing the platform as ML work expands

An ML platform rarely arrives fully formed; it grows with the organisation's machine-learning work. Many organisations start with a single GPU server for an initial team or project, and as more teams take up ML, more projects run, and demand rises, the platform expands — adding GPUs to the pool, storage for growing datasets and models, and the scheduling to share a larger pool among more users. Planning a platform means allowing for this growth, so that it can start at a sensible size and scale as the organisation's ML work does, rather than being sized once and outgrown or over-provisioned from the start.

This makes a platform something to size to current needs with room to grow, rather than a fixed build. A first dedicated GPU server can be the seed of a platform, serving early ML work while the practices around sharing, environments, and storage take shape; as demand grows, the pool and its supporting infrastructure expand to match. Hosting that supports this growth — letting capacity be added as needed — lets a platform track the organisation's ML maturity. When we host an ML platform, we size it to where the organisation is and plan for where it is heading, so that it can begin proportionate to current work and grow into a larger shared foundation as the ML effort expands, without a disruptive rebuild at each step.

Sovereignty for the whole organisation's ML, and where we fit

Because an ML platform holds and processes the organisation's ML work across teams and projects, its sovereignty is the sovereignty of the organisation's ML data as a whole. Datasets, models, prompts, and the data flowing through the platform's pipelines all sit on it, so where the platform runs governs all of that. VV Internet Hosting is incorporated in the Netherlands, within the EU, so a self-hosted ML platform hosted with us runs under European jurisdiction and outside the direct reach of the US CLOUD Act. For an organisation whose ML data is sensitive or subject to European protection, hosting the whole platform in the EU keeps all of that data under European law — a comprehensive sovereignty that self-hosting on EU infrastructure provides and a managed foreign platform cannot.

We host the dedicated infrastructure for a self-hosted ML platform: a shared pool of GPUs, common environments, shared storage, and a place for the scheduling and orchestration that tie them together — dedicated to your organisation, in EU datacenters under European jurisdiction. This suits organisations consolidating their ML work onto a platform they control, where the shared pool is economical and where the sovereignty of the organisation's ML data matters. We are candid about our limits: if you want a fully managed platform with a provider's tooling and no infrastructure to operate, that is a different kind of provider. We give you the sovereign, dedicated foundation to run your own platform on — the right choice for organisations that want control and European data sovereignty across their ML work, and not for those seeking a fully managed service.

Questions

ML platform, answered plainly

Common questions about hosting for ML platform.

What is an ML platform, versus a single pipeline?

An ML platform is the shared foundation an organisation's teams run their machine-learning work on — a common pool of compute, environments, storage, and orchestration serving many projects and pipelines. A pipeline is one workflow; a platform is the layer beneath many pipelines and projects, giving them a common, managed base rather than each building its own.

Why does an ML platform improve cost efficiency?

By pooling expensive GPUs so they're shared across teams rather than dedicated per team. Dedicated per-team GPUs sit idle whenever that team isn't using them; a shared pool is used by whoever has demand, keeping it busier. Higher utilisation of expensive hardware directly lowers the cost of the organisation's ML compute, so the total pool can be smaller.

Should we self-host an ML platform or use a managed one?

It depends on what you value. A managed platform is convenient but places your ML work on the provider's infrastructure. A self-hosted platform on infrastructure you control keeps the compute, data, and terms in your hands, at the cost of operating it. Organisations often self-host when their ML data becomes too sensitive for a provider-run platform.

Planning ML platform infrastructure?

We host dedicated, EU-sovereign infrastructure sized to your workload — and we will tell you plainly when something else fits better. Tell us what you're building.