actaserve

From power
to useful work.

Choosing compute means matching a whole system to a workload. Follow the dependencies from the electrical supply to the cost of a task that actually meets your needs.

How a model generates a response
In this guide

Start with the measurement boundary.

A chip’s power rating, a workstation’s wall draw and a facility’s electricity bill describe different things. Before comparing numbers, draw a boundary around what the meter includes.

Accelerator boardHost, memory, network and PSU lossesFacility cooling and electrical overhead

Watts describe power at a point in time or averaged over a stated interval. Joules describe energy over time. A short burst at high power can use less energy than a long task at lower power; neither a board rating nor idle draw tells you which workload completes more efficiently.

Energy (joules) = integral of power (watts) over time (seconds)

Use metered energy across the same workload window. Include failed attempts and retries when the goal is energy per accepted task.

Compare like measurement boundaries.

Board power
A board rating excludes the host, other system components and facility overhead. It is not a measured energy cost per request.
System wall power
State which components and peripherals the meter includes, the workload, utilisation and measurement interval. A maximum draw or PSU capacity rating is not typical serving power.
Facility energy
Include cooling and electrical overhead within a stated boundary and period. Do not compare a facility total directly with a board rating.

Unmatched ratings cannot support an efficiency ranking. That requires matched work, quality criteria and power boundaries.

The room is part of the system.

The electrical path runs through supply, distribution and power conversion before reaching the processors. Capacity planning must account for equipment limits, sustained loading, redundancy and service procedures—not just the sum of accelerator ratings.

Heat has to leave the building.

Air cooling moves heat from components into an airflow path. A liquid-cooled component still needs a way to reject that heat, whether through a room-side heat exchanger or a facility loop. “Liquid cooled” does not, on its own, specify the building requirements.

Confirm inlet conditions, airflow direction, rack density, cooling connections where applicable, noise limits, floor loading, access clearances and maintenance space. A compact system suitable for one room may need a different operating arrangement in another.

Use PUE carefully.

PUE = total facility energy ÷ IT equipment energy

The numerator and denominator must cover the same period and facility boundary.

PUE describes facility overhead, not model efficiency. Multiplying IT energy by a representative PUE can estimate facility energy, but an annual site average is not a measured marginal overhead for one job. Do not multiply again if the starting meter already includes facility consumption.

Specify what happens during a power or cooling fault: which systems stop, what jobs can resume, and who responds. Resilience has both a capacity cost and an operating cost.

A chip is not a serving system.

Chip: execute operationsSystem: host, memory, storage, networkCluster: distribute and schedule work

The host prepares inputs, schedules work, handles networking and runs application or tool logic. Accelerators execute supported model operations. This is a division of work, not a universal rule that CPUs cannot run inference or accelerators only generate tokens.

Software determines usable hardware.

Compatible API interfaces are a starting point, not a promise of identical model behaviour. Check the endpoints and capabilities your application uses. For your own model, confirm the checkpoint, licence and serving requirements before committing to capacity.

Check the checkpoint, model architecture, precision, supported operators, compiler and serving runtime together. Nominal capacity is not proof that a particular model loads, runs correctly or meets the response target. Reserve memory for runtime buffers and request state as well as weights.

Storage affects model distribution and initial loading. Once loaded, weights normally stay resident while the runtime serves many requests. A warm service does not reload the entire model from storage for each prompt.

Scaling out introduces movement.

Data parallelism places replicas on separate devices and distributes requests. Model parallelism partitions one model’s work across devices. A partitioned model must exchange intermediate data; the placement and interconnect become part of the latency and throughput question.

Adding devices adds potential capacity, not a guaranteed linear speedup. Memory locality, synchronisation, communication, scheduler behaviour and request lengths can limit the result. Separate the network carrying user requests from the links carrying model-parallel traffic in the deployment design.

Compare memory architectures.

Memory capacity answers what can fit. Bandwidth describes how quickly data can move under a particular access pattern. Latency describes the delay for an access or operation. They are related constraints, not interchangeable scores.

Unified memory: shared CPU / GPU capacity.

In a unified-memory architecture, CPU and GPU share memory rather than separate host and device pools. Sharing avoids a separate CPU-to-GPU copy for data that can remain in the shared allocation; it does not eliminate internal data movement, memory pressure or runtime overhead.

Weights, KV caches, temporary buffers and other system users compete for available capacity. More simultaneous long contexts can constrain serving even when the model weights fit.

Unified memory Step-by-step walkthrough
One shared addressable memory architectureStep through the architecture
  1. 1. Share

    CPU and GPU access unified memory rather than two separate host and device pools.

    CPUPrepare work
    Unified memoryModel + context
    GPURun model
    Shared capacityWeights + context + system work
  2. 2. Read

    The GPU reads resident weights and request state from that same memory. Sharing a pool does not remove bandwidth limits.

    CPURuntime
    Unified memoryWeights + KV cache
    GPURead for inference
    Unified-memory bandwidthShared access still has bandwidth limits
  3. 3. Make room

    More concurrent contexts need more memory alongside weights, runtime buffers and the operating system. Installed memory is not all free for model weights.

    Model weightsKV cachesRuntime + OS
    Capacity trade-offLarger contexts leave less room for concurrency
Memory topology; areas and links are schematic, not a speed comparison.

Many-core: local memory, distributed work.

A programmable many-core accelerator distributes operations across cores and local memories. In a system with separate accelerator memory, the runtime must place working data where the operations need it.

Summed local bandwidth across chips is not the bandwidth of a single shared memory pool, nor the bandwidth of the interconnect. A partitioned workload has to map its working data to those local memories and exchange data where its operations require it.

Many-core distributed memory Step-by-step walkthrough
Local memory does not become one poolStep through the architecture
  1. 1. One card

    In this schematic, the accelerator has its own local memory. The runtime maps supported model operations onto its cores.

    HostRuntime
    AcceleratorLocal memory
    Local bandwidthData must reach the cores
  2. 2. Partition work

    This schematic partitions work across chips with separate local memory. Model partitioning requires data exchange; the chip count is illustrative.

    Chip 1Local memoryChip 2Local memoryChip 3Local memoryChip 4Local memory
    Exchange between model partitions
    Partitioned capacitySeparate local memories
  3. 3. Keep host separate

    Host and accelerator memory are separate in this example. Summed local bandwidth is not bandwidth available to one shared allocation.

    AcceleratorLocal memoryAcceleratorLocal memoryAcceleratorLocal memoryAcceleratorLocal memory
    Host memory / separate pool
    Aggregate local bandwidthNot one shared memory pool
Memory topology; areas and links are schematic, not a speed comparison.

Dataflow: map operations and data movement.

Reconfigurable dataflow architectures let a compiler map operations and their data paths onto compute resources. Memory organisation varies by design; there is no universal set of memory tiers.

Check where working data resides and how it reaches each operation. The compiler and runtime determine the mapping and movement. Confirm model support, local capacity and access costs rather than interpreting all memory bytes as equally accessible.

Dataflow mapping Step-by-step walkthrough
Map operations and working dataStep through the architecture
  1. 1. Locate working data

    Identify the compute resources and memory available in the proposed design. Local storage and access paths differ between architectures.

    Compute resourcesWorking dataMemory access paths
    Design boundaryMemory organisation varies
  2. 2. Stage working data

    The compiler maps operations and data paths; the runtime manages execution. Placement determines where working data must move.

    Mapped operationsWorking dataData paths
    Data movementPlacement shapes access costs
  3. 3. Check the mapping

    Whether a model serves well depends on supported operations, working-set placement and data movement, not just total memory capacity.

    Model supportWorking-set fitAccess costs
    System fitCheck model mapping, not just total bytes
Memory topology; areas and links are schematic, not a speed comparison.

Follow inference and application workflows.

Initialisation is separate from inference.

  1. Load and retain weights. The runtime initialises the model before requests are served. Switching or evicting models can require loading again, but this is not the normal per-request loop.
  2. Admit and prefill a request. A scheduler may queue the request. Prefill processes its prompt positions and creates attention keys and values: the KV cache. These are request context, not new learned model parameters.
  3. Select the first token. The final prompt position’s output scores support selection of the first generated token. Time to first token includes waiting and prompt processing as well as the response path.
  4. Decode repeatedly. Each iteration processes the previously generated token, reads cached context, appends new keys and values and selects the next token. Generated tokens become part of the context as they are processed. Weights remain resident and unchanged.
  5. End and reclaim. A stop token, output limit or cancellation ends the request. Request memory can be reclaimed; prefix caching and retention policies may keep reusable state under runtime-specific rules.

A longer prompt increases prefill work and cache size. A longer answer adds decode iterations. More concurrent requests add cache demand. Batching may improve aggregate throughput while changing per-request waiting time. Some runtimes use sliding windows, cache compression or speculative decoding; the walkthrough demonstrates the basic autoregressive dependency rather than every optimisation.

How a model generates a response Step-by-step walkthrough
From resident weights to a responseStep through the architecture
  1. 1. Initialise once

    The runtime loads model weights into memory before serving. They stay resident across requests while the model remains loaded.

    InputNo request
    Resident modelWeights retained
    OutputResponse
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    No request cache
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative inputWaiting for a request
  2. 2. Send a prompt

    A request brings new input, not a new copy of the weights. It can wait in a queue until the scheduler admits it.

    InputWhy does memory matter?
    Resident modelWeights retained
    OutputResponse
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    No request cache
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative inputWhy does memory matter?
  3. 3. Prefill

    The model processes the prompt. Attention keys and values for its positions form this request’s KV cache, so later tokens can use that context.

    InputProcess prompt positions
    Resident modelWeights retained
    OutputResponse
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    Prompt context
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative inputPrompt positions processed
  4. 4. First token

    Prefill produces the scores used to select the first output token. The response starts; model weights are unchanged.

    InputPrompt consumed
    Resident modelWeights retained
    OutputNext token
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    Prompt context
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative outputMemory
  5. 5. Decode

    Feed the previous output token back through the resident model. Read the existing context, append its new keys and values, then select the next token.

    InputPrevious token: Memory
    Resident modelWeights retained
    OutputNext token
    Previous output → resident model + cached context → next output
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    Prompt contextOutput 1
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative outputMemory matters
  6. 6. Decode again

    Repeat the same decode operation. Output extends one token at a time and the request’s KV cache grows as generated tokens are processed.

    InputPrevious token: matters
    Resident modelWeights retained
    OutputNext token
    Previous output → resident model + cached context → next output
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    Prompt contextOutput 1Output 2
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative outputMemory matters.
  7. 7. Finish the request

    A stop condition ends generation. The runtime can reclaim this request’s cache; the model stays loaded for the next request.

    InputRequest ended
    Resident modelWeights retained
    OutputResponse
    Model memoryWeights residentInitialisation is not repeated per request.
    Request memory / KV cache
    No request cache
    Context state, not learned weights. Schematic; blocks are not to scale.
    Illustrative outputMemory matters.
Autoregressive text inference. Scheduling, batching and cache retention vary by runtime. These steps explain dependencies, not elapsed time.

Application latency is bigger than model latency.

Coding agents
The model returns a tool request; the harness checks authority, executes the tool, adds its result to context and calls the model again. Tool execution and dependent calls contribute to task latency. A fast token rate alone does not establish a fast or correct coding task.
Coding agents: context, tools and results Step-by-step walkthrough
Agent / tool loopStep through the architecture
  1. 1. Send context

    The harness sends instructions and relevant repository context to the model.

    HarnessContext + permissions
    Context envelope
    instructionsselected files
    Model APIText + tool calls
    Tool environmentRead / run / return
    HarnessContinue or finish
    Harness sendsinstructions + selected files
    Tool result returns to the harness, then the next model call.
  2. 2. Request a tool

    The model returns a structured tool call. It does not run a shell or read your filesystem by itself.

    HarnessContext + permissions
    Model APIText + tool calls
    Tool call envelope
    tool namearguments
    Tool environmentRead / run / return
    HarnessContinue or finish
    Model requestsread_file(path)
    Tool result returns to the harness, then the next model call.
  3. 3. Execute outside the model

    The harness checks permissions and runs the requested tool in its execution environment.

    HarnessContext + permissions
    Model APIText + tool calls
    Tool environmentRead / run / return
    Tool execution
    permission checkfile read
    HarnessContinue or finish
    Harness executesauthorised file read
    Tool result returns to the harness, then the next model call.
  4. 4. Return the result

    The harness adds the tool result to the conversation and calls the model again. This next model call depends on that result.

    HarnessContext + permissions
    Expanded context
    instructionsselected filestool result
    Model APIText + tool calls
    Tool environmentRead / run / return
    HarnessContinue or finish
    Tool result joins contextfile contents → next model call
    Tool result returns to the harness, then the next model call.
  5. 5. Continue or answer

    The model uses the returned context to choose another tool call or a response. The harness decides when the task is done.

    HarnessContext + permissions
    Model APIText + tool calls
    Tool environmentRead / run / return
    HarnessContinue or finish
    Model decision
    next tool callor response
    Model returnsnext action or proposed answer
    Tool result returns to the harness, then the next model call.
Agent/tool sequence. Each dependent step waits for the one before it.
Voice
External speech recognition produces text, the text-model API generates response text, and external speech synthesis produces audio. The voice application owns turn detection, buffering and interruption cancellation. Measure first-audio latency and stale-response handling across that full chain.
Voice: speech services around a text model Step-by-step walkthrough
Conversation across three servicesStep through the architecture
  1. 1. Recognise speech

    A separately supplied speech-recognition service turns an utterance into text.

    Speech recognitionExternal service
    Audio input
    audio chunkaudio chunkaudio chunk
    Text modelActaserve API
    Speech synthesisExternal service
    Voice applicationPlayback + turn-taking
    External speech recognitionaudio → transcript
    Turn-taking and cancellation belong to the voice application.
  2. 2. Generate text

    Your voice application sends that transcript and conversation context to the text-model API.

    Speech recognitionExternal service
    Text modelActaserve API
    Text input
    transcriptconversation context
    Speech synthesisExternal service
    Voice applicationPlayback + turn-taking
    Actaserve text-model roletranscript → response text
    Turn-taking and cancellation belong to the voice application.
  3. 3. Synthesise audio

    A separately supplied synthesis service turns response text into audio. First-audio latency includes the stages before it.

    Speech recognitionExternal service
    Text modelActaserve API
    Speech synthesisExternal service
    Audio output
    audio chunkaudio chunkaudio chunk
    Voice applicationPlayback + turn-taking
    External speech synthesisresponse text → audio
    Turn-taking and cancellation belong to the voice application.
  4. 4. Handle an interruption

    When the user speaks again, your voice stack stops or cancels outstanding output and updates the conversation before the next turn.

    Speech recognitionExternal service
    Text modelActaserve API
    Speech synthesisExternal service
    Voice applicationPlayback + turn-taking
    New turn
    old output cancellednew utterance
    Voice application controlscancel old response / begin new turn
    Turn-taking and cancellation belong to the voice application.
Conversational pipeline: separately supplied speech components around a text-model API.
Video generation
Submit through POST /v1/videos/generations, keep the job ID and poll GET /v1/videos/generations/:id. Handle pending, completed and failed outcomes. Polling is not a new generation request; a retry policy must avoid accidentally creating duplicate jobs.
Video: submit, track and retrieve a job Step-by-step walkthrough
An asynchronous job, not a streamStep through the architecture
  1. 1. Submit

    Your application submits a prompt and supported generation settings. It must handle submission errors before tracking a job.

    ApplicationPrompt + settings
    Submission
    promptsettings
    Generation APIJob identifier
    Supplier jobPending / processing
    ApplicationResult or error
    Submit requestPOST /v1/videos/generations
  2. 2. Keep the job ID

    The API returns an identifier. Your application keeps it so the user can leave the submission screen while the job is pending.

    ApplicationPrompt + settings
    Generation APIJob identifier
    Job record
    job ID retainedpending
    Supplier jobPending / processing
    ApplicationResult or error
    Application storesjob ID
  3. 3. Poll status

    Use the same job ID to check status while the supplier processes the job. Polling observes the job; it does not submit another generation.

    ApplicationPrompt + settings
    Generation APIJob identifier
    Supplier jobPending / processing
    Same job record
    job ID retainedstatus check
    ApplicationResult or error
    Check statusGET /v1/videos/generations/:id
  4. 4. Handle the result

    When complete, use the returned result reference. A failed job needs an error state, not a result preview; pending work stays pending.

    ApplicationPrompt + settings
    Generation APIJob identifier
    Supplier jobPending / processing
    ApplicationResult or error
    Terminal branch
    completed: result referencefailed: error
    Branch on terminal statuscompleted → result / failed → error
Asynchronous video generation: submit a prompt, track the job, retrieve the result.
Robotics and sensing
Sensor observations become perception estimates. Local placement can reduce dependence on a remote link, but safety control and actuation authority need a separate specification. A network-dependent model is not a safety interlock. This is a deployment-design discussion, not a deployed fleet service.
Robotics: perception and separate safety control Step-by-step walkthrough
Perception and control have different jobsStep through the architecture
  1. 1. Capture observations

    Sensors produce observations. Decide the sample rate and what can be discarded before moving a continuous stream.

    SensorsObservations
    Observations
    sensor framerange samples
    Local perceptionEstimates, not authority
    Remote analysisOptional network path
    Sensor inputframes / range observations
    Separate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.
  2. 2. Run local perception

    A perception model converts observations into estimates. Keeping it near the sensor changes the dependence on network connectivity.

    SensorsObservations
    Local perceptionEstimates, not authority
    Estimates
    objectspositionsuncertainty
    Remote analysisOptional network path
    Perception outputobjects / positions / uncertainty
    Separate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.
  3. 3. Choose what travels

    Selected observations or summaries may go to remote analysis when the link allows. Network loss must be part of the deployment design.

    SensorsObservations
    Local perceptionEstimates, not authority
    Remote analysisOptional network path
    Selected data
    summaryoptional remote analysis
    Optional remote workloadselected data → analysis
    Separate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.
  4. 4. Keep safety separate

    Safety logic must independently constrain actuation. A perception estimate or remote model response is not permission to move.

    SensorsObservations
    Local perceptionEstimates, not authority
    Remote analysisOptional network path
    Separate control authoritysafety inputs → interlocks → actuation
    Separate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.
    Independent control
    safety inputsinterlocksactuator limits
Perception architecture. Keep safety control within a separate, independently specified boundary.

Actaserve supports coding/text/tool integration and asynchronous video generation through the current API surfaces, subject to model and supplier support. Speech services and robotics safety-control software are separately supplied.

Price the useful result, not the headline rate.

Start with a workload definition: model and licence, precision, prompt and output lengths, traffic shape, concurrency, tool use and acceptance criteria. Keep the incumbent and candidate on the same test contract. There is no matched Actaserve coding benchmark in this guide.

Keep the denominators visible.

Time to first token
Request start to first output. State whether the measurement includes client networking and queueing, and report the distribution rather than only its average.
Output tokens per second
Generation rate after output begins. Label per-request versus aggregate throughput. Neither describes prompt processing time or tool delays.
Accepted tasks per period
Completed tasks that pass the agreed quality and safety criteria. Include retries and failures in the cost numerator even when they do not count as accepted output.
Cost per accepted task = total attributable cost ÷ accepted tasksEnergy per accepted task = measured joules ÷ accepted tasks

If no tasks pass, cost per accepted task is undefined—not zero.

Compare complete operating arrangements.

For an API, account for billed usage, retries and tool or storage services outside the model charge. For dedicated capacity, include the reservation or hardware cost over its useful life, finance where relevant, energy, facility costs, network, software, staffing, support, spares and downtime. Avoid counting the same cost in both a bundled service fee and a separate line item.

Utilisation spreads fixed costs over useful work. At a fixed period cost, halving accepted task volume doubles fixed cost per accepted task. That arithmetic is not a performance forecast: raising utilisation can also increase queues and miss latency targets. Leave headroom for bursts, failures and maintenance.

Report demand assumptions, idle periods, acceptance rate, depreciation period, residual value and what the meter excludes. Evaluate a range of demand levels instead of treating a single full-utilisation estimate as the business case.

Run acceptance before committing capacity.

  • Pin the checkpoint, runtime version, precision and sampling settings.
  • Use representative task inputs with authorised data and explicit pass criteria.
  • Record warm and cold conditions separately, including concurrency and request-length distributions.
  • Measure first response, complete-task latency, quality, failures, memory and power over a matched window.
  • Repeat under expected sustained traffic and bursts. Record what happens during supplier, network and host failures.

Only measured results from that configuration can establish whether it meets your workload’s targets. Manufacturer specifications establish equipment characteristics, not service outcomes.

Put sovereignty in the operating terms.

A hardware brand does not establish data residency, isolation or ownership. Specify processing, storage, backups and permitted transfers; administrator and root access; support access; tenant isolation; maintenance and incident response; and the model and data handover process.

Dedicated compute reservations are open. Delivery is scoped per project: Actaserve will own and operate the hardware, including procurement, colocation power/cooling, model bring-up and ongoing operations. Customers will choose root/bare-metal access or a managed API endpoint. Private-weight deployments will be custom quoted, subject to model, licence and runtime qualification.

API access does not reserve a particular chip. Dedicated hardware availability, location, capacity, price, lead time and access terms are agreed per project. The architectures here are illustrative, not an availability guarantee.

Bring the workload definition and the boundaries you need to protect. The configuration discussion should end in explicit delivery and acceptance terms, not an assumed hardware guarantee.

Discuss your compute requirements

Sources and reading notes.

This guide explains general architecture mechanisms and measurement boundaries, not a particular product’s specifications or measured performance. Hardware examples, manufacturer figures and manufacturer citations are not part of this public guide.

The step-driven explanation approach was informed by MTS Compute. Its models, calculators and performance data are not Actaserve evidence. Here, weight loading is explicitly separated from per-request prefill and decode.