Nearby Computing
Request a demo

Resources  |  Blog

Introducing AI Mesh: deploy, connect and govern AI, from one platform

I’m really excited to write this post and finally disclose what has kept us busy at Nearby Computing during 2026. For several months we have been building a module for the problems our customers keep running into when they operate AI: where a model may run, how it reaches the rest of the estate, who may call it and what it costs. Along the way, NearbyOne gained AgentOne, an assistant agent built into the product, so our customers get help from our own agent as they work; we joined NVIDIA Inception, NVIDIA’s programme for startups; and we shaped the product around the challenges of distributed inference, which we take role by role further down.

We are launching it today as Nearby AI Mesh, a NearbyOne module for deploying, connecting and governing AI workloads across every cloud, edge site and GPU factory an organisation runs. AI Mesh extends the NearbyOne control plane with what enterprise AI has been missing: a continuously reconciled service fabric that spans clouds, per-tenant isolation at layer 7, and one governed checkpoint, the AI gateway, that every AI call passes through and where every token is counted. It ships today as a technical preview.

We are a European company, and NearbyOne, with AI Mesh now part of it, is built in Europe for organisations that must operate distributed inference under the EU AI Act, NIS2 and the Sovereignty Effectiveness Assurance Levels (SEAL). The platform installs on infrastructure our customers own and does not call home, so models and data stay where they were placed, and the estate can show where every model runs, which project may reach it and what leaves the perimeter. NearbyOne is built to help our customers meet those requirements and show that they meet them.

NearbyOne already covers the rest of the estate: Nearby Glide, our Cloud Manager, runs workloads across clouds and edge sites; Nearby Beam manages networks, from private 5G to SD-WAN and the other elements that connect sites to the edge; and Nearby Forge, our Infrastructure Manager, runs servers and devices as one IT fleet. AI Mesh complements them from the angle none of them covers: governing AI. This post explains why we built it, through four real-world challenges we see at our customers, then how it works under the hood and how it is governed, how it meets each of the four challenges, how it fills the roles of a distributed inference architecture, and where we believe the discipline of AI orchestration is heading.

Part 1 of 5 · AI Mesh in practice

This is the first in a series of posts on Nearby AI Mesh. Here we use four examples to illustrate what customers need today. The following posts take each need in turn and go into the technical detail: scaling out to providers and back, AI as a service on a shared fleet, keeping AI inside a perimeter, and knowing what it costs.

The AI Dashboard in the AI Estate Manager. A row of tiles for AI entities on their sites and gateways, health, the GPU fleet, tokens over 24 hours and the reconciliation of intents; then panels for the cloud sites and their GPUs, the AI projects, what needs attention, the models through the AI gateway with input and output tokens, egress and sovereignty showing what stays on the platform and what leaves the perimeter, and MCP access.
Figure 1: The AI Dashboard in the AI Estate Manager, the portal inside AI Mesh, with the whole AI estate on one page.

AI has become a distributed systems problem

Two years ago, a typical enterprise AI deployment was one model behind one endpoint on one cloud. That world is gone. The applications our customers run today are pipelines: agents calling Model Context Protocol (MCP) servers, retrieval services grounding prompts in enterprise data, routers steering requests between self-hosted models and managed providers. And classical software is becoming agentic, which multiplies the number of AI components that must be deployed, secured and governed.

At the same time, the infrastructure underneath has fragmented. Inference runs where the data, the latency budget or the regulator says it must: a retail back-room server, a telco edge site, a private OpenShift cluster, a hyperscaler region, a GPU factory. Sovereignty rules pin data and compute in-country. GPU and egress costs climb unpredictably.

The result is that enterprise AI is no longer primarily a modelling problem. It is a distributed systems problem: placement, connectivity, identity, isolation, observability and cost attribution across a heterogeneous estate. The industry has spent a decade building this machinery for microservices. Almost none of it, as deployed today, spans clouds, and very little of it understands tokens, models or GPUs.

That gap is what we mean by AI orchestration, and it is the gap AI Mesh closes.

Real-world challenges

AI Mesh grew out of what we kept seeing at customers. Their sectors differ, but the same few challenges come back: GPUs that run out at the worst moment, one fleet that has to serve many customers without any of them seeing each other, AI that must never leave a perimeter, and spend that nobody can put a number on. Four examples, one per challenge, show them best.

You’re a retailer that runs an open-weight model on its own GPUs for store assistants and back-office agents, and on a sales weekend those GPUs fill up. And you keep asking yourself: could the workload offload itself when they are full, and come back on its own when the peak is over, without anyone touching the application?

You’re a telco that runs its own audio processing, such as call transcription and summarisation, on its own GPUs, and your business customers now want models and GPU capacity from you too. And you keep asking yourself: how could I sell models as a service and GPUs as a service from one fleet, keep my own workloads on it, and still see which customer used what?

You serve a defence customer that accepts coding assistants and document drafting only from models running on its own premises, so your agents and developers must stay inside the perimeter. And you keep asking yourself: if one of them reached a model outside, would I even know?

You’re a bank that has many teams using models, some on shared GPUs and some from outside providers. Finance sees the provider invoices a month late, and the models you run yourselves send none at all. And you keep asking yourself: what is AI actually costing us, and who is spending it?

Each of these challenges can be met today, and we come back to all four once we have explained how AI Mesh works.

What AI Mesh addresses

Different as they sound, those four challenges come down to three problems that came up again and again in customer environments, and AI Mesh is built to address all three. They are usually run as separate projects by separate teams; we think they are one job.

  1. AI across multiclouds and technologies
    Teams need to run a mix of AI components, models, agents, MCP servers, vector stores, on whatever substrate each site provides: Rancher or OpenShift, CPU or GPU, AMD, Arm, Intel or NVIDIA silicon. And they sit behind whichever AI gateway each team picked, Agent Router (formerly Envoy AI Gateway), agentgateway or LiteLLM, each needing its own routes, keys and fallback. Most stitch this together by hand, cluster by cluster and gateway by gateway.
  2. Multi-tenancy
    Isolation and discovery are usually a trade-off. Lock the estate down and services cannot find each other; open it up and tenants leak into each other. Service providers hosting many customers on shared GPU infrastructure feel this most acutely.
  3. Governance, cost and compliance
    Nobody can answer the simple questions: who may call which model, where does the data go, and what is our AI actually costing? Self-hosted inference produces no invoice at all. Managed-provider spend hides in a dozen consoles. And every service tends to carry its own authentication, its own rate limits and its own blind spots.

How AI Mesh works

AI Mesh applies NearbyOne’s core architectural idea, declarative intent plus continuous reconciliation, to the AI estate. Operators declare which workloads, which connectivity and which policies apply to which tenant; the platform renders the required Kubernetes and gateway resources into every cluster and runs a reconciliation loop that detects and corrects drift. The same declarative profile produces a consistent deployment on Rancher or OpenShift, on CPU nodes or GPU nodes.

NearbyOne control plane Deploy · connect · govern, from declared intent, continuously reconciled YOUR ESTATE On premises Your data centres Edge Close to the data Public cloud Your cloud regions AI factory GPU model serving AI gateway AI gateway AI gateway AI gateway Governance at every gateway: access · limits · sovereignty · metering AI Service Fabric: joins the gateways across sites, one private plane per tenant External AI APIs Managed model providers Outside the perimeter NearbyOne control plane Deploy · connect · govern YOUR ESTATE On premises Your data centres Edge Close to the data Public cloud Your cloud regions AI factory GPU model serving AI gateway AI gateway AI gateway AI gateway Governance at every gateway: access · limits · sovereignty · metering AI Service Fabric: joins the gateways across sites, one plane per tenant External AI APIs Managed model providers Outside the perimeter
Figure 2: How AI Mesh is put together. One NearbyOne control plane deploys, connects and governs AI wherever it runs: on premises, at the edge, in public cloud and in GPU AI factories. An AI gateway at each site applies the same governance, the AI Service Fabric joins the gateways across sites, and a call to an external AI API leaves the perimeter only through a gateway.

Architecturally, the module is organised around four layers.

The substrate: dissimilar machines, one platform. NearbyOne onboards servers (including zero-touch provisioning at the edge), installs and lifecycle-manages Kubernetes, and renders the infrastructure operators that make specialised hardware usable, for example the NVIDIA GPU Operator for drivers, device plugins and MIG partitioning. A fleet of dissimilar machines behaves as one platform rather than a collection of snowflakes. The telemetry stack is rendered into every cluster as part of the same profile, not bolted on afterwards: kube-state-metrics and DCGM for object and GPU state, Prometheus for the pipeline and Thanos for long-term, cross-cluster retention.

The catalogue: vetted building blocks for any site. A curated marketplace of versioned, packaged blocks: agents, MCP servers, models and model platforms, vetted by an organisation’s own AI, compliance and infrastructure teams. Model serving can run on open-source stacks, such as vLLM, Ollama and KServe, or on NVIDIA NIM and NVIDIA Triton Inference Server, and managed providers sit beside them, all deployable to any cloud site with one click or via API.

The service fabric: connecting every site, isolating every tenant. It is what makes the rest of AI Mesh possible. The AI Service Fabric carries traffic directly from one Kubernetes cluster to another at layer 7, with no public load balancers or provider virtual IPs in the path, and lands it natively on the destination’s data plane. One name resolves everywhere, so a service is addressed the same way from every site. The plane runs east-west across CPU clouds and edge sites as well as north-south into the AI factory: a tenant’s agents in cloud A, its retrieval-augmented generation (RAG) index in cloud B and its large language model (LLM) endpoints in the GPU factory sit on one private plane and address each other as if local. Declare that service X must reach service Y, and the service fabric keeps a working path alive as endpoints scale, move or fail over.

The multi-tenancy model falls out of the same mechanism. The service fabric publishes each tenant’s services into that tenant’s own namespaces, and the AI gateways live in those same namespaces. Isolation and reachability are two sides of one namespace-scoped plane: Tenant A’s components find each other by name across every cloud, while remaining invisible and unreachable to Tenant B.

On premises Data centre Public cloud Cloud region AI factory GPU site Tenant A its own plane Agent Knowledge base LLM No route, no name, no address between the two planes Tenant B its own plane Chat assistant MCP server LLM Tenant A its own plane Tenant B its own plane On premises Data centre Agent Chat assistant Public cloud Cloud region Knowledge base MCP server AI factory GPU site LLM LLM No route, no name, no address between the two planes
Figure 3: One plane per tenant. Both tenants run components on the same three sites, yet each tenant's agents, knowledge bases and models reach each other by name on that tenant's plane only. Neither tenant can see or address the other.

The gateway layer: one front door for every model call. At the edges of the service fabric sit multi-tier AI gateways, built on state-of-the-art open-source components that evolve every day. A Tier-1 gateway at the aggregation site is the organisation’s one front door; Tier-2 gateways front the self-hosted serving stacks, such as KServe and vLLM, in each AI factory. The Tier-1 gateway routes each request, applies the project’s rate limits and meters every token. The gateway presents one OpenAI-compatible API in front of self-hosted models and managed providers alike, translating each provider’s native API and keeping the credentials in one place, so a team can switch or mix providers without touching application code.

Deploying a model through NearbyOne wires it into the gateway, and into the gateway’s fallback order, without anyone writing a route. That holds whichever AI gateway the customer runs, including the three named earlier: NearbyOne configures the gateway, so the choice of gateway stays the customer’s.

Because every model call, inbound to a self-hosted model or outbound to a managed provider, passes through the gateway, it is the natural and uniform point for policy and metering. That single design decision is what makes the governance, and the economics, work.

Automation: you compose and activate, AI Mesh does the rest. Teams compose their service and activate it. AI Mesh carries out the rest: every configuration, every interconnect between sites and every deployment, and keeps each of them as declared.

Governance: one checkpoint, policies that follow the workload

Connectivity and cost are means to an end. What customers ask for first is control: who may call which model, with what data, where it may run, and what happens when a limit is reached. AI Mesh answers that at the one point every AI call already passes through.

Because NearbyOne deploys the workloads itself, it knows the site, the jurisdiction and the GPU behind every endpoint. AI gateways are connected automatically as models are deployed, so a project gets one endpoint that reaches its own models on premises and at the edge, a public cloud region, and external AI APIs.

Gateways that connect themselves let an AI workload stretch from your own GPUs to the edge, the public cloud and external AI APIs, behind one endpoint. Policy decides how far it goes.

That policy has four parts:

  • Access. Each project is given its models, agents and MCP servers. Each tenant’s services sit on a plane no other tenant can reach, and the MCP access matrix shows which consumer may use which tool server.
  • Limits. Request and token rates per project, budgets with alert thresholds, and quotas at organisation, project and site level. A limit you set stops the call; it does not send it somewhere else.
  • Sovereignty. Every site carries a region in the Sovereignty Engine, and everything on it inherits that region. Placement decides jurisdiction, not the kind of model, and every call that leaves the perimeter is visible, typed by what crosses: inference, embedding, retrieval or control.
  • Evidence. A live map, a bill of materials and an intent-versus-reconciled view, read from what the platform deployed rather than from a spreadsheet.

The AI Estate Manager, the portal inside AI Mesh, draws all of this as one picture: the AI Map. Each organisation has one front door, its Tier-1 AI gateway. Every site’s gateway is wired to it automatically as models are deployed, and consumers such as chat interfaces and agents reach every model through that front door, never by addressing a site directly. Because jurisdiction comes from where a thing is placed, the map shows at a glance what stays inside a jurisdiction and what crosses to an external provider. The map’s layers switch between entities, infrastructure and jurisdictions, with sovereignty as an overlay. The retailer’s answer further down is this picture at work: a model in a Tier-2 AI factory and the same model at an external provider, both behind one front door.

EUROPEAN UNION AI factory Tier-2 Models AI gateway AI factory Tier-2 Models AI gateway SPAIN (EU) Headquarters Tier-1 Tier-1 AI gateway the one front door Chat interface Agent Knowledge base stays on-platform MCP server OUTSIDE THE PERIMETER External provider Managed model publishes to Tier-1 uses a model through the front door uses MCP leaves the perimeter SPAIN (EU) Headquarters Tier-1 Knowledge base stays on-platform MCP server Chat interface Agent Tier-1 AI gateway the one front door AI factory Tier-2 AI gateway Models AI factory Tier-2 AI gateway Models EUROPEAN UNION External provider Managed model OUTSIDE publishes to Tier-1 uses a model through the front door uses MCP leaves the perimeter
Figure 4: The AI Map, as a concept. Consumers at headquarters reach every model through one Tier-1 front door; each AI factory's gateway publishes its models to that front door, wired automatically as models are deployed. Jurisdiction comes from where each site is, so the one call that leaves the perimeter, to a managed model at an external provider, is drawn as a crossing.

Meeting the four challenges

With the pieces in place, here is how AI Mesh meets each of the four challenges today, in the same order.

The retailer: scaling out to a provider, and back

The retailer runs gpt-oss, an open-weight model, on its own GPUs in a multi-GPU AI factory. Its store assistants and back-office agents all call one unified endpoint: one address, with as many models behind it as the team wants, where each request asks for the model it needs. A model behind that endpoint can be served on premises, by a provider, or both.

For gpt-oss the team lists two instances, in order. First, the model on premises, on the multi-GPU site. Second, the same model from a public provider: OpenAI’s gpt-oss 20B. Each instance is tried only when the one above it cannot answer, because it is unreachable, because it is failing, or because it turns the request away. The team puts them in that order in the AI Estate Manager, and NearbyOne configures the AI gateway’s priority routing to match: local first, provider second. Both sit behind the retailer’s one Tier-1 front door, so applications call the same address either way.

The Which models answer screen of a unified endpoint in the AI Estate Manager. A Qwen failover model is served on premises on the Multinode-multiGPU site. gpt-oss has two ordered instances: 1st, on premises on Multinode-multiGPU; 2nd, from the provider OpenAI as gpt-oss 20B, with arrows to reorder them.
Figure 5: Which models answer, in the AI Estate Manager. One endpoint with several models behind it; gpt-oss has two instances in order, first on premises on the multi-GPU site and second from OpenAI as gpt-oss 20B. Each instance is tried when the one above it cannot answer.

Most days the first instance answers everything. Then comes a peak, say a sales weekend, and demand runs past what the local GPUs can serve. The local model starts turning requests away, and the AI gateway does what its fallback order says: each request the local model cannot take goes on to the provider, through the same endpoint, with no change to any application. When demand falls, requests are answered locally again on their own, because the on-premises instance is always tried first.

The trigger is not a forecast. The gateway does not watch GPU utilisation and guess; it moves a request on only when the instance above cannot answer it. A limit is different: a limit the team sets stops the call instead of sending it on, so a rate cap never quietly becomes a bill somewhere else.

Requests to one endpoint What your own GPUs can answer 1 Your own GPUs answer every request 2 The local model turns the excess away; the gateway sends it to the provider 3 Demand falls: every request answered on premises again Morning Peak Evening 1st: on premises, tokens priced at your rate card 2nd: public provider, tokens priced at its rate Requests to one endpoint 1 2 3 Morning Peak Evening 1 Your own GPUs answer every request. 2 Past capacity, the local model turns the excess away and the gateway sends it to the provider, through the same endpoint. 3 As demand falls, every request is answered on premises again. 1st: on premises, your rate card 2nd: public provider, its own rate What your own GPUs can answer
Figure 6: Cascading models over a day (illustrative). The local model answers everything up to what the on-premises GPUs can serve. At the peak it turns the excess away and the gateway sends that overflow to the public provider through the same endpoint; as demand falls, every request is answered on premises again. Tokens are counted on both paths.

Every token is counted on both paths. The provider’s tokens are priced at the provider’s rate and the local ones at the retailer’s own rate card, so the cost of the peak is one figure in one place, not a surprise on a provider invoice a month later. And the AI Map shows the provider call for what it is: a crossing out of the perimeter, typed as inference.

This is the stretching in practice. Nobody redeployed anything or wrote a route. The model was deployed, the gateway was wired with local first, and policy decided how far the workload could reach.

Part 2 We go further in Scale out to a provider, and come back on your own, coming soon in this series.

The telco: models and GPUs as a service on one fleet

The telco gives each business customer a tenant of its own, with its own plane and its own projects, and each project carries its own rate limits and quotas. The telco’s own transcription and summarisation run as projects on the same fleet, and no customer can see or reach another’s services, even where they share a site. Models, and the GPU capacity to run them, are ordered from one marketplace and deployed on the telco’s own sites, so models as a service and GPUs as a service come from the same catalogue.

To see who used what, the gateway meters every token per tenant, and the GPU fleet and consumption views show what each pool of GPUs is doing and which models ran on it. Two things are not there yet. Turning that usage into each customer’s bill is the per-project chargeback that follows as it lands, and letting a business customer resell what it bought, with limits nested inside the telco’s, is on the list at the end.

One marketplace: models as a service and GPUs as a service THE TELCO BUSINESS CUSTOMERS, EACH ON ITS OWN PLANE Audio processing Transcription Summarisation Tokens used Customer A Its models Its limits Tokens used Customer B Its models Its limits Tokens used Customer C Its models Its limits Tokens used One GPU fleet, on your own sites every token metered per tenant at the gateway One marketplace: models and GPUs as a service THE TELCO Audio processing Transcription Summarisation Tokens used BUSINESS CUSTOMERS, EACH ON ITS OWN PLANE Customer A Its models Its limits Tokens used Customer B Its models Its limits Tokens used Customer C Its models Its limits Tokens used One GPU fleet, on your own sites every token metered per tenant
Figure 7: One fleet, many customers. The telco's own transcription and summarisation and each business customer's models run on the same GPU fleet; every customer has its own plane and its own limits, and its tokens are metered at the gateway. Models and GPU capacity come from one marketplace.

Part 3 We go further in Models and GPUs as a service, from one fleet, coming soon in this series.

The defence customer: every model on premises

The answer here is placement. The customer’s models are deployed onto its own on-premises site, and the project’s endpoint lists only those instances, so there is nowhere else for a request to go. Its agents, MCP servers and knowledge bases sit on its own plane.

Would you know if something reached outside? The Egress and sovereignty view shows everything in the estate that reaches outside the perimeter, consumer by consumer and typed by what crosses, so the evidence is a list you expect to be empty, and you can see at once when it is not. That is something you can see today, not something that is stopped: refusing a call by jurisdiction, an approval catalogue that decides the allowed sites, and a per-call audit record are all on the list at the end. The view reads from what the platform deployed, so a tool that goes around the platform altogether is a matter for the network.

YOUR PERIMETER, ON PREMISES Developers coding assistants Agents document drafting Project endpoint on-premises models only On-prem model On-prem model Egress and sovereignty Leaves the perimeter Nothing: the list is empty OUTSIDE External provider No route to it from your endpoint Managed model YOUR PERIMETER, ON PREMISES Developers coding assistants Agents document drafting Project endpoint on-premises models only On-prem model On-prem model Egress and sovereignty Leaves the perimeter Nothing: the list is empty It shows a crossing if one appears. OUTSIDE External provider Managed model No route to it from your endpoint
Figure 8: Every model on premises. Coding assistants and agents call a project endpoint whose models all run inside the perimeter, so there is no route from it to an external provider, and the Egress and sovereignty view lists nothing leaving. The view shows a crossing if one appears; it does not block it.

Part 4 We go further in Every model at home, and the evidence to show it, coming soon in this series.

The bank: what AI costs, and who spends it

Self-hosted inference produces no invoice, so the bank sets the price of its own tokens: a rate card, applied to the tokens its own models serve. Calls to outside providers are priced at each provider’s published rate. Every token is metered on every path, and both kinds land in one estimated figure for the organisation as the calls happen, not a month later.

That answers what AI is costing. Who is spending it starts with the projects: each team gets one with a budget, an alert threshold and a quota, so AI spending has an owner before it has a bill. Today the cost figure is showback for the organisation as a whole; it splits into money per team once per-project chargeback lands.

Tokens on premises × Your rate card Tokens from providers × Provider rate AI spend one estimated figure for the organisation, as the calls happen EACH TEAM: A PROJECT WITH A BUDGET, AN ALERT THRESHOLD AND A QUOTA Team A Team B alert Team C alert threshold Bars show usage against each team's quota. Cost per team follows with per-project chargeback. Tokens on premises × Your rate card Tokens from providers × Provider rate AI spend one estimated figure for the organisation, as the calls happen EACH TEAM: A PROJECT WITH A BUDGET, AN ALERT THRESHOLD AND A QUOTA Team A Team B alert Team C alert threshold Bars show usage against each team's quota. Cost per team follows with per-project chargeback.
Figure 9: What AI costs, and who spends it. Tokens served on premises are priced at the bank's own rate card and provider tokens at the provider's rate, giving one estimated figure for the organisation. Each team's project carries a budget, an alert threshold and a quota; the bars show usage against each team's quota, and cost per team follows with per-project chargeback.

Part 5 We go further in What AI costs, and who is spending it, coming soon in this series.

What distributed inference needs, and how AI Mesh fills it

A distributed inference architecture spreads GPU capacity over many sites: an organisation’s own data centres, the edge, public clouds and AI factories. Over that capacity sits a control plane that decides what runs where, who may reach it, and how each request finds a model. AI Mesh is built to be that control plane, on the GPUs an organisation already runs.

Control-plane roles in a distributed inference architecture REFERENCE ROLE FILLED IN AI MESH BY Orchestrator NearbyOne: cluster, GPU and model lifecycle Resource and Intent Engine Declared intent, continuously reconciled Decision and Policy Engine Tenant planes, projects and residency rules LLM router and AI gateway AI gateways per tenant, in fallback order Global load balancing Tier-1 gateway, the one entry point Orchestrator dashboard AI Estate Manager PROVIDED BY THE INFRASTRUCTURE GPUs and accelerators what the models run on Datacentre network fabric the switches and DPUs between GPUs Control-plane roles in a distributed inference architecture, and what fills them ORCHESTRATOR NearbyOne: cluster, GPU and model lifecycle RESOURCE AND INTENT ENGINE Declared intent, continuously reconciled DECISION AND POLICY ENGINE Tenant planes, projects and residency rules LLM ROUTER AND AI GATEWAY AI gateways per tenant, in fallback order GLOBAL LOAD BALANCING Tier-1 gateway, the one entry point ORCHESTRATOR DASHBOARD AI Estate Manager PROVIDED BY THE INFRASTRUCTURE GPUs and accelerators what the models run on Datacentre network fabric the switches and DPUs between GPUs
Figure 10: What distributed inference needs, and how AI Mesh fills it. The control-plane roles of a distributed inference architecture on the left, and what fills each in AI Mesh on the right. The GPUs, the accelerators and the network fabric between them are provided by the infrastructure; AI Mesh starts from a working cluster with working GPUs.

Role by role:

  • Orchestrator. NearbyOne handles the lifecycle of clusters, GPUs and models on every site, with Nearby Forge for the machines and Nearby Glide for the workloads.
  • Resource and Intent Engine. Declared intent, continuously reconciled: the engine every NearbyOne module shares.
  • Decision and Policy Engine. Per-organisation isolation, access and residency rules: the per-tenant plane, projects with their limits, and the Sovereignty Engine.
  • LLM router and AI gateway. The AI gateways, one set per tenant, route each request to a model and fall back in the order set for it. We claim the role, not every policy a router could run: prompt-aware or KV-cache-aware routing is not something AI Mesh does today.
  • Global load balancing. The Tier-1 aggregation gateway, the organisation’s one entry point.
  • Orchestrator dashboard. The AI Estate Manager.

AI Mesh runs on NVIDIA GPUs, with the GPU Operator and DCGM, and on AMD, Arm and Intel. What it does not do is build or watch the network inside the AI factory. The datacentre fabric that joins the GPUs, with its switches and DPUs, belongs to the infrastructure and to the partners who supply it, and what it does is theirs to describe; AI Mesh starts from a working cluster with working GPUs.

What we deliberately did not build

Credibility in a new category comes from being clear about boundaries. AI Mesh is a control plane, not a cloud: it does not provide the GPUs, replace your Kubernetes distributions or host the models. It operates across what you already run, and it assumes teams comfortable with cloud-native operations. We also chose open foundations, the Kubernetes Gateway API, vLLM, KServe and Prometheus, and we work with the AI gateways customers already use, precisely so that adopting AI Mesh never means betting your architecture on us. Orchestration should remove lock-in, not add a new layer of it.

Where AI orchestration goes next

Our view is that within a few years, “where should this model run and who may reach it” will be as unremarkable a question as “where should this container run” is today, answered by policy, not by meetings. The organisations that get there first will be the ones that treat connectivity, isolation, governance and cost as properties of one service fabric rather than four separate projects.

That is the direction we are taking AI Mesh. None of what follows is a switch you can turn on today; it is where the work is going:

  • Keys bound to a project and a person, so every call through the gateway is attributable to both.
  • Budgets that act by falling back to your own model instead of stopping the call.
  • Jurisdiction-aware routing that fails closed: a request that may not leave its jurisdiction is refused, not sent.
  • An approval catalogue that also decides the allowed sites, so approving a model also says where it may run.
  • A per-call audit record for the EU AI Act and NIS2.
  • Nested limits, so a telco can resell GPU capacity to its own customers and let them set limits for theirs.

If you want to see it against your own estate, we suggest a scoped pilot: one tenant, one on-prem cluster, one public provider, with the service fabric, the gateway fallback and the AI Dashboard live in days.

The series

This is the first of five posts. The next four take one technical challenge each, illustrated here by one of our four examples, and go into how AI Mesh meets it:

  • Part 1Introducing AI Meshthe overview
    You are here
  • Part 2Scale out to a provider, and come back on your ownfallback across your GPUs and providers
    Coming soon
  • Part 3Models and GPUs as a service, from one fleetmulti-tenancy on shared GPUs
    Coming soon
  • Part 4Every model at home, and the evidence to show itsovereignty and the perimeter
    Coming soon
  • Part 5What AI costs, and who is spending ittokenomics and budgets
    Coming soon

Request a pilot Explore AI Mesh

Written by

David Carrera

CTO and Co-Founder of Nearby Computing. David is the chief architect of NearbyOne and leads product strategy and technical direction.

Follow us for new articles

We publish every few weeks on LinkedIn and here on the blog.

Follow on LinkedIn