AI compute.
On your terms.

Build with our model API. Reserve dedicated compute with GPU or alternative architectures, qualified for your workload.

Conceptual sculpture of open, connected compute: white modular layers joined by a lime-green ribbon.
Read context
Generate
Return answer
Context in. Tokens out.

A familiar way to build.
More room to choose.

OpenAI-compatible APIAnthropic-compatible APIExplore the docs

Start building.
Make room for more.

From a first model call to a dedicated deployment, choose the access that fits your workload.

Usage-based pricing

Model API

Bring your application. Choose a model. Connect through the interfaces your team already knows.

  • OpenAI- and Anthropic-compatible
  • Models from the active catalogue
  • Published rates for each model
Get an API key

Model capabilities and pricing apply. API access does not reserve dedicated hardware or guarantee data residency.

Reservations open

Dedicated compute

A deployment shaped around your models, your traffic and where the work needs to run.

  • Root/bare-metal access or a managed API endpoint
  • Hardware procurement and ongoing operations
  • Bring private weights, subject to qualification
Reserve capacity

Custom quote. Availability, location, access rights and delivery terms are scoped per project. Deployment terms.

Explore pricing

One request.
Different kinds of work.

A dedicated deployment can use GPUs to read the prompt and inference silicon to generate the response. The model and workload decide the fit.

See how inference works

Placement map

Illustrative deployment

RequestCoding agent

Long repository context

Read the promptPrefill

GPU

KV cacheTransfer between machines
Generate the responseDecode

Inference ASIC

ResponseYour application

Streamed tokens

/v1/chat/completions
Follow one request from context to response.

Coding agent: GPUs could read the repository context, then transfer its KV cache to an inference ASIC to generate the reply. Model and runtime qualification determine whether this split is possible.

Illustrative placement, not live telemetry. Each split is qualified per model before it carries traffic; API access does not select this hardware.

What will you build?

Connect through the API, or explore the compute behind your next application.

Give your agent
room to think.

API integration

Use text and tool calls through the API, including the Pi coding-harness integration. Your harness controls tool permissions and execution.

Explore the agent workflow
ContextInstructions + selected filesModel callActaserve APITool resultYour authorised harness
Each model call depends on the context and tool results before it. Your harness supplies the next turn.

Keep the
conversation moving.

Integration architecture

Use the text-model API between separately supplied speech recognition and synthesis components. Your application manages turn-taking and interruptions.

Explore the voice chain
Speech to textExternal serviceText modelActaserve APIText to speechExternal service
Delay adds up across the chain. Your application cancels stale output when the speaker interrupts.

From a prompt
to a video job.

Generation API

Submit a prompt, keep the job ID and check for completion. Available models depend on the active supplier catalogue.

Explore a video job
Submit promptStart a generation jobKeep the job IDCheck the same jobReceive resultOr handle a failed job
GET /v1/videos/generations/:id. Asynchronous generation, separate from live video understanding.

Bring compute
closer to the work.

Deployment design

Discuss local perception, memory and power for sensing workloads, alongside separately supplied robotics software and safety control.

Explore the perception boundary
SensorsFrames + observationsLocal perceptionEstimates, not authorityRemote analysisOptional network path
Safety control stays separate. No direct actuation path from the model; network loss is part of deployment design.

Different silicon.
More ways forward.

The model, memory and traffic shape the choice. Explore the architectures behind a dedicated deployment.

GPU

Prefill & batch throughput

Thousands of parallel cores beside dedicated high-bandwidth memory. A strong fit where the work is compute-bound.

See prefill and decode

Dataflow

Latency-sensitive decode

The compiler maps operations and data movement onto reconfigurable compute resources. Memory organisation varies by design.

Explore dataflow mapping

Many-core

Parallel throughput

Work is distributed across cores and local memories. Model mapping and data movement shape the fit.

Explore distributed memory

Unified memory

Large footprint, fewer users

CPU and GPU share a memory pool. Weights, context and system work compete for that capacity.

Explore shared memory

Illustrative architectures, not an availability guarantee. Memory architectures / Power and cooling

Your workload.
Our groundwork.

We'll own and operate your dedicated hardware. You choose root/bare-metal access or a managed API endpoint.

Reserve capacity

Hardware

We buy and own the machines for your contract, then qualify the model and serving stack.

Allocation

Capacity is assigned to your contract, with the delivery date set in your quote.

Sites

Power, cooling and location are scoped around the hardware and agreed deployment terms.

Operations

Model bring-up, serving and monitoring for the life of the contract.

What do you
have in mind?

Tell us what you need to run. We'll come back with an architecture, pricing and a delivery window.

hello@actaserve.com

Include your models, expected traffic or concurrency, response-time target, and location or access requirements. Share estimates where you can.

Don't include prompts, credentials or customer data. We use Formspree to receive your enquiry. Privacy notice.