dot point
Platform Updates

‍hosted·ai v2.5: token factories, multi-GPU, and the most advanced GPU scheduler yet

September 8, 2026

We’re excited to announce the availability of hosted·ai v2.5, with a brand new GPU scheduler and new GPU revenue opportunities for neoclouds - now including token factories. Let’s get into it!

‍

‍

What's new in hosted·ai v2.5:

‍

Scheduler v2 - bringing fluidity to GPU infrastructure

Users get better performance and uptime. Providers get better utilization, more scalability, and better manageability for hosting production inference workloads.

‍

The headline feature in this release is a re-engineered GPU scheduler. The scheduler is the heart of the hosted·ai platform: it turns GPU into an elastic compute resource, solves the problem of idle GPU, and enables neoclouds to make more per GPU hour while delivering lower-cost services for customers.

In the existing scheduler, we provisioned multi-tenant workloads to a pool of GPUs, but each workload was tied to a specific physical card.

In hosted·ai v2.5 the scheduler completely decouples AI workloads from physical GPUs and brings a new level of fluidity to multi-tenant GPU infrastructure. It assigns workloads to any GPU in the pool and continuously load-balances workloads across GPUs, to make optimal use of the resources available.

  • Dynamic scheduling maximizes utilization of GPU resources in the pool, enabling more efficient hosting of multi-tenant inference workloads
  • The scheduler automatically migrates workloads to the most suitable GPU available to optimize performance (if GPUs are becoming overloaded) and to ensure uptime (should a GPU fail, for example)
  • It controls and polices resource access and thresholds: tenants have no direct access to the GPU hardware, improving security and manageability
  • Live migration of workloads also simplifies neocloud operations and SLOs for production inference - in future releases this will be extended to cross-node migration

‍

Multi-GPU support - up to eight GPUs per workload

‍Host more demanding AI models on each node: multiple tenants can provision up to eight GPUs per workload.

‍

With the new scheduler, up to eight virtualized GPUs can now be assigned to a single workload. This feature works in tandem with hosted·ai’s multi-tenant scheduling capabilities, so that multiple users can map multiple GPUs to workloads simultaneously.

Neocloud providers can now host more demanding AI models on each node. Multi-GPU support makes use of the NVIDIA Collective Communication Library (NCCL).

‍

GPU emulation - rent out GPUs you didn’t buy

Now customers don’t have to find a new vendor because you didn't have the GPU they wanted.

‍

Customers don’t always need the latest and greatest GPU as a service. They don’t always want to pay a premium for cutting-edge GPU models.

In hosted·ai v2.5, you can present a pool of modern GPUs as any previous generation of card. Now you can serve different customer price points and use cases without buying a large mixed GPU fleet.

  • For example, a pool of NVIDIA GB300s could simultaneously be presented to customers as A100s, H100s or B200s, as well as GB300s
  • Customers pay for the equivalent resources of the smaller card, and you control that pricing
  • Resources are not reserved: they are provisioned on demand by the hosted·ai scheduler

‍

Token Factory - become an inference cloud

Build your own token factory. Host OpenAI-compatible open-weight models, and turn GPUs into token revenue.

‍

Token Factories are available as a new service type in hosted·ai v2.5, alongside bare metal, GPUaaS based on Kubernetes, and VM passthrough based on KVM.

  • Now you can set up public or private token factory services and sell hosted endpoints for a huge range of open-source LLMs, priced per million tokens.
  • Token Factories are fully compatible with the OpenAI API: customers just change a line of code to switch to your service.
  • Token Factories can be built with on-premises GPU or GPU from our capacity trading network (GPU Mesh).‍

You can see the first live deployment of a token factory at packet.ai.

‍

Advanced networking

Private VLAN, VXLAN and IP improvements. More flexible enterprise networking.

‍

Also new in hosted·v2.5: private VLAN capabilities, enabling tunnelling for neocloud customers who need to connect to another data center or cloud environment; VXLAN overlay networks, connecting virtualized GPU instances in a region on a private address range; and dedicated IPs with web service mapping, instead of shared endpoint port redirection, now IPs follow the workload when it moves.

‍

‍Next steps

This release also includes many smaller features and fixes based on customer feedback.

  • To upgrade from previous versions to hosted·ai v2.5, please contact your account manager or our customer success team.
    ‍
  • If you're new to hosted·ai, get in touch for a demo and we'll walk/talk you through the platform. Thanks!

‍