Equinix Inference Exchange is moving enterprise AI infrastructure closer to users and data, combining Equinix’s global data-center footprint with Nvidia reference architectures and Together AI’s open-model inference platform.

 

Announced on September 2, the service is designed for companies that want production inference across multiple locations without stitching together compute, networking, cooling and model-serving systems on their own. Equinix says availability will begin in the first quarter of 2027.

 

How Equinix Inference Exchange Is Built

The architecture divides responsibility across three partners. Equinix supplies the physical infrastructure, power, cooling, operations and Equinix Fabric connectivity. Nvidia provides validated Enterprise Reference Architectures for accelerated AI infrastructure, while Together AI supplies the inference layer and access to more than 200 open-source models.

 

That combination targets a practical bottleneck in enterprise AI: models may be easy to test through an API, but running them reliably in production requires enough compute capacity, predictable network paths, operational support and a location strategy that matches where applications and regulated data actually live.

 

The service combines four operational layers:

  • Equinix data centers with power, advanced cooling and day-two infrastructure operations
  • Nvidia Enterprise Reference Architectures for production AI compute
  • Together AI multitenant or dedicated single-tenant inference environments
  • Equinix Fabric links to clouds, networks, data sources and other AI providers

 

Equinix says its broader footprint spans more than 280 data centers in 77 metros, with 230 cloud on-ramps and more than 10,500 businesses interconnected across its exchange. The company also says eight of the top 10 AI model providers and nine of the top 10 AI clouds already have deployments in its ecosystem.

 

The service was announced alongside Equinix Fabric One, a separate connectivity offering intended to automate how enterprises link distributed cloud, network and AI environments. Together, the announcements show Equinix treating interconnection as part of the AI infrastructure stack rather than a separate networking task.

 

Why Distributed Inference Changes Enterprise AI

Training large models is concentrated in a relatively small number of huge clusters, but inference has a different geography. A bank, retailer, manufacturer or software company may need responses generated near customers in several regions while keeping sensitive data inside particular jurisdictions.

 

Equinix is pitching Inference Exchange around that distributed pattern. Metro-edge deployments could reduce the distance between an application and the inference system serving it. Sovereign AI deployments could place processing in locations chosen for residency requirements, while open-model migration could let companies move selected workloads away from proprietary APIs.

 

The design also reflects a shift from buying AI hardware as an isolated project toward treating compute, networking and operations as one system. Inference performance can be limited by more than GPU speed; network latency, model placement, capacity contention and data movement can all affect the experience seen by an end user.

 

Related Coverage

 

What Together AI Adds to the Stack

Together AI is important because the service is not being framed around a single flagship model. Equinix says the platform will support more than 200 open-source models, giving customers a way to select models by workload rather than locking the infrastructure to one provider’s API.

 

The partners plan both shared multitenant deployments and dedicated single-tenant environments. The shared model can spread infrastructure costs across users, while dedicated capacity is aimed at organizations that need stronger isolation or more predictable access to compute.

 

For Equinix, that flexibility turns interconnection into part of the AI product. A company could keep inference near a private dataset, connect the deployment to public clouds or internal networks through Equinix Fabric, and still change the model layer without rebuilding every surrounding connection.

 

The Q1 2027 Rollout Still Has Open Questions

The announcement sets a launch window but leaves several procurement details unresolved. Equinix has not yet published initial deployment metros, pricing, detailed hardware configurations, capacity allocations or specific service-level performance commitments.

 

Those omissions matter because geography is central to the pitch. A sovereign-AI buyer needs to know exactly which jurisdictions are available, while a latency-sensitive application needs metro-level placement and network-performance information before comparing the service with a hyperscaler or dedicated GPU cloud.

 

The Q1 2027 launch will therefore be measured less by the size of Equinix’s global footprint than by how much of that footprint becomes usable for inference on day one. Customer wins, published pricing and concrete metro availability will show whether Inference Exchange becomes a broad enterprise platform or begins as a narrower set of regional deployments.