AI compute.On your terms.
Build with our model API. Reserve dedicated compute with GPU or alternative architectures, qualified for your workload.
Context in. Tokens out. Play sequence
A familiar way to build.More room to choose.
Start building. Make room for more. From a first model call to a dedicated deployment, choose the access that fits your workload.
{ } Usage-based pricing
Model API Bring your application. Choose a model. Connect through the interfaces your team already knows.
OpenAI- and Anthropic-compatible Models from the active catalogue Published rates for each model
Get an API key ↗
Model capabilities and pricing apply. API access does not reserve dedicated hardware or guarantee data residency.
[ ] Reservations open
Dedicated compute A deployment shaped around your models, your traffic and where the work needs to run.
Root/bare-metal access or a managed API endpoint Hardware procurement and ongoing operations Bring private weights, subject to qualification
Reserve capacity ↗
Custom quote. Availability, location, access rights and delivery terms are scoped per project. Deployment terms .
Explore pricing →
One request. Different kinds of work. A dedicated deployment can use GPUs to read the prompt and inference silicon to generate the response. The model and workload decide the fit.
See how inference works ↗
Placement map Illustrative deployment
Coding agent
Voice turn
Document batch
Private model
Request Coding agent Long repository context
Read the prompt Prefill
GPU
KV cache Transfer between machines
Generate the response Decode
Inference ASIC
Response Your application Streamed tokens
/v1/chat/completions
Follow one request from context to response. Play request
Coding agent: GPUs could read the repository context, then transfer its KV cache to an inference ASIC to generate the reply. Model and runtime qualification determine whether this split is possible.
Illustrative placement, not live telemetry. Each split is qualified per model before it carries traffic; API access does not select this hardware.
What will you build? Connect through the API, or explore the compute behind your next application.
Coding agents Voice Video Robotics & sensing
Give your agent room to think. API integration
Use text and tool calls through the API, including the Pi coding-harness integration. Your harness controls tool permissions and execution.
Explore the agent workflow ↗
ContextInstructions + selected files Model callActaserve API Tool resultYour authorised harness
Each model call depends on the context and tool results before it. Your harness supplies the next turn.
Keep the conversation moving. Integration architecture
Use the text-model API between separately supplied speech recognition and synthesis components. Your application manages turn-taking and interruptions.
Explore the voice chain ↗
Speech to textExternal service Text modelActaserve API Text to speechExternal service
Delay adds up across the chain. Your application cancels stale output when the speaker interrupts.
From a prompt to a video job. Generation API
Submit a prompt, keep the job ID and check for completion. Available models depend on the active supplier catalogue.
Explore a video job ↗
Submit promptStart a generation job Keep the job IDCheck the same job Receive resultOr handle a failed job
GET /v1/videos/generations/:id. Asynchronous generation, separate from live video understanding.
Bring compute closer to the work. Deployment design
Discuss local perception, memory and power for sensing workloads, alongside separately supplied robotics software and safety control.
Explore the perception boundary ↗
SensorsFrames + observations Local perceptionEstimates, not authority Remote analysisOptional network path
Safety control stays separate. No direct actuation path from the model; network loss is part of deployment design.
Different silicon. More ways forward. The model, memory and traffic shape the choice. Explore the architectures behind a dedicated deployment.
GPU Prefill & batch throughput + Thousands of parallel cores beside dedicated high-bandwidth memory. A strong fit where the work is compute-bound.
See prefill and decode ↗
Dataflow Latency-sensitive decode + The compiler maps operations and data movement onto reconfigurable compute resources. Memory organisation varies by design.
Explore dataflow mapping ↗
Many-core Parallel throughput +
Unified memory Large footprint, fewer users + CPU and GPU share a memory pool. Weights, context and system work compete for that capacity.
Explore shared memory ↗
Illustrative architectures, not an availability guarantee. Memory architectures / Power and cooling
Your workload. Our groundwork. We'll own and operate your dedicated hardware. You choose root/bare-metal access or a managed API endpoint.
Reserve capacity ↗
Hardware We buy and own the machines for your contract, then qualify the model and serving stack.
Allocation Capacity is assigned to your contract, with the delivery date set in your quote.
Sites Power, cooling and location are scoped around the hardware and agreed deployment terms.
Operations Model bring-up, serving and monitoring for the life of the contract.
What do you have in mind? Tell us what you need to run. We'll come back with an architecture, pricing and a delivery window.
hello@actaserve.com ↗