Tenancy and isolation
Detail — shared against dedicated“Dedicated” is a claim about where your work physically runs. On a shared endpoint your request is batched onto a GPU alongside other customers' requests because that is what makes the economics work. We do not need to do that: we own the accelerators, so we can hold them for one tenant and still price the way sheet 02 describes.
Shared replica
Four tenants batched onto one GPU — four hatches, one machine.
Dedicated replica
One tenant inside an isolation boundary. Yours for the life of the reservation.
Tenancy
A replica holds one customer's work. Your requests are never batched into a shared queue with anyone else's, because nothing else is resident on the GPU.
Training
Your prompts, documents, sequences and code are never used to train or fine-tune anything, ours or anyone's.
Retention
Inputs and outputs are held only for the life of the request, or the life of an agent run where the run needs its own transcript. Nothing persists past it unless you ask us to keep it.
Isolation
Each customer's replicas, storage and keys are separated. An agent run executes in its own sandbox with no route to another tenant's data.
Keys
Scoped per project, revocable, and never printed back to you after issue.
Placement
Workloads are pinned to named halls in a facility we operate, so you can be told which machines ran your work.
These are architectural commitments, not certifications. Nerin holds no third-party compliance attestation yet; ask us where that stands before moving a regulated workload.