Start with the measurement boundary.
A chip’s power rating, a workstation’s wall draw and a facility’s electricity bill describe different things. Before comparing numbers, draw a boundary around what the meter includes.
Watts describe power at a point in time or averaged over a stated interval. Joules describe energy over time. A short burst at high power can use less energy than a long task at lower power; neither a board rating nor idle draw tells you which workload completes more efficiently.
Energy (joules) = integral of power (watts) over time (seconds)Use metered energy across the same workload window. Include failed attempts and retries when the goal is energy per accepted task.
Compare like measurement boundaries.
- Board power
- A board rating excludes the host, other system components and facility overhead. It is not a measured energy cost per request.
- System wall power
- State which components and peripherals the meter includes, the workload, utilisation and measurement interval. A maximum draw or PSU capacity rating is not typical serving power.
- Facility energy
- Include cooling and electrical overhead within a stated boundary and period. Do not compare a facility total directly with a board rating.
Unmatched ratings cannot support an efficiency ranking. That requires matched work, quality criteria and power boundaries.
The room is part of the system.
The electrical path runs through supply, distribution and power conversion before reaching the processors. Capacity planning must account for equipment limits, sustained loading, redundancy and service procedures—not just the sum of accelerator ratings.
Heat has to leave the building.
Air cooling moves heat from components into an airflow path. A liquid-cooled component still needs a way to reject that heat, whether through a room-side heat exchanger or a facility loop. “Liquid cooled” does not, on its own, specify the building requirements.
Confirm inlet conditions, airflow direction, rack density, cooling connections where applicable, noise limits, floor loading, access clearances and maintenance space. A compact system suitable for one room may need a different operating arrangement in another.
Use PUE carefully.
PUE = total facility energy ÷ IT equipment energyThe numerator and denominator must cover the same period and facility boundary.
PUE describes facility overhead, not model efficiency. Multiplying IT energy by a representative PUE can estimate facility energy, but an annual site average is not a measured marginal overhead for one job. Do not multiply again if the starting meter already includes facility consumption.
Specify what happens during a power or cooling fault: which systems stop, what jobs can resume, and who responds. Resilience has both a capacity cost and an operating cost.
A chip is not a serving system.
The host prepares inputs, schedules work, handles networking and runs application or tool logic. Accelerators execute supported model operations. This is a division of work, not a universal rule that CPUs cannot run inference or accelerators only generate tokens.
Software determines usable hardware.
Compatible API interfaces are a starting point, not a promise of identical model behaviour. Check the endpoints and capabilities your application uses. For your own model, confirm the checkpoint, licence and serving requirements before committing to capacity.
Check the checkpoint, model architecture, precision, supported operators, compiler and serving runtime together. Nominal capacity is not proof that a particular model loads, runs correctly or meets the response target. Reserve memory for runtime buffers and request state as well as weights.
Storage affects model distribution and initial loading. Once loaded, weights normally stay resident while the runtime serves many requests. A warm service does not reload the entire model from storage for each prompt.
Scaling out introduces movement.
Data parallelism places replicas on separate devices and distributes requests. Model parallelism partitions one model’s work across devices. A partitioned model must exchange intermediate data; the placement and interconnect become part of the latency and throughput question.
Adding devices adds potential capacity, not a guaranteed linear speedup. Memory locality, synchronisation, communication, scheduler behaviour and request lengths can limit the result. Separate the network carrying user requests from the links carrying model-parallel traffic in the deployment design.
Compare memory architectures.
Memory capacity answers what can fit. Bandwidth describes how quickly data can move under a particular access pattern. Latency describes the delay for an access or operation. They are related constraints, not interchangeable scores.
Unified memory: shared CPU / GPU capacity.
In a unified-memory architecture, CPU and GPU share memory rather than separate host and device pools. Sharing avoids a separate CPU-to-GPU copy for data that can remain in the shared allocation; it does not eliminate internal data movement, memory pressure or runtime overhead.
Weights, KV caches, temporary buffers and other system users compete for available capacity. More simultaneous long contexts can constrain serving even when the model weights fit.
Unified memory Step-by-step walkthrough
Choose a step
Reduced motion is on. Choose each step manually.
1. Share
CPU and GPU access unified memory rather than two separate host and device pools.
CPUPrepare workUnified memoryModel + contextGPURun modelShared capacityWeights + context + system work2. Read
The GPU reads resident weights and request state from that same memory. Sharing a pool does not remove bandwidth limits.
CPURuntimeUnified memoryWeights + KV cacheGPURead for inferenceUnified-memory bandwidthShared access still has bandwidth limits3. Make room
More concurrent contexts need more memory alongside weights, runtime buffers and the operating system. Installed memory is not all free for model weights.
Model weightsKV cachesRuntime + OSCapacity trade-offLarger contexts leave less room for concurrency
Many-core: local memory, distributed work.
A programmable many-core accelerator distributes operations across cores and local memories. In a system with separate accelerator memory, the runtime must place working data where the operations need it.
Summed local bandwidth across chips is not the bandwidth of a single shared memory pool, nor the bandwidth of the interconnect. A partitioned workload has to map its working data to those local memories and exchange data where its operations require it.
Many-core distributed memory Step-by-step walkthrough
Choose a step
Reduced motion is on. Choose each step manually.
1. One card
In this schematic, the accelerator has its own local memory. The runtime maps supported model operations onto its cores.
HostRuntimeAcceleratorLocal memoryLocal bandwidthData must reach the cores2. Partition work
This schematic partitions work across chips with separate local memory. Model partitioning requires data exchange; the chip count is illustrative.
Chip 1Local memoryChip 2Local memoryChip 3Local memoryChip 4Local memoryExchange between model partitionsPartitioned capacitySeparate local memories3. Keep host separate
Host and accelerator memory are separate in this example. Summed local bandwidth is not bandwidth available to one shared allocation.
AcceleratorLocal memoryAcceleratorLocal memoryAcceleratorLocal memoryAcceleratorLocal memoryHost memory / separate poolAggregate local bandwidthNot one shared memory pool
Dataflow: map operations and data movement.
Reconfigurable dataflow architectures let a compiler map operations and their data paths onto compute resources. Memory organisation varies by design; there is no universal set of memory tiers.
Check where working data resides and how it reaches each operation. The compiler and runtime determine the mapping and movement. Confirm model support, local capacity and access costs rather than interpreting all memory bytes as equally accessible.
Dataflow mapping Step-by-step walkthrough
Choose a step
Reduced motion is on. Choose each step manually.
1. Locate working data
Identify the compute resources and memory available in the proposed design. Local storage and access paths differ between architectures.
Compute resourcesWorking dataMemory access pathsDesign boundaryMemory organisation varies2. Stage working data
The compiler maps operations and data paths; the runtime manages execution. Placement determines where working data must move.
Mapped operationsWorking dataData pathsData movementPlacement shapes access costs3. Check the mapping
Whether a model serves well depends on supported operations, working-set placement and data movement, not just total memory capacity.
Model supportWorking-set fitAccess costsSystem fitCheck model mapping, not just total bytes
Follow inference and application workflows.
Initialisation is separate from inference.
- Load and retain weights. The runtime initialises the model before requests are served. Switching or evicting models can require loading again, but this is not the normal per-request loop.
- Admit and prefill a request. A scheduler may queue the request. Prefill processes its prompt positions and creates attention keys and values: the KV cache. These are request context, not new learned model parameters.
- Select the first token. The final prompt position’s output scores support selection of the first generated token. Time to first token includes waiting and prompt processing as well as the response path.
- Decode repeatedly. Each iteration processes the previously generated token, reads cached context, appends new keys and values and selects the next token. Generated tokens become part of the context as they are processed. Weights remain resident and unchanged.
- End and reclaim. A stop token, output limit or cancellation ends the request. Request memory can be reclaimed; prefix caching and retention policies may keep reusable state under runtime-specific rules.
A longer prompt increases prefill work and cache size. A longer answer adds decode iterations. More concurrent requests add cache demand. Batching may improve aggregate throughput while changing per-request waiting time. Some runtimes use sliding windows, cache compression or speculative decoding; the walkthrough demonstrates the basic autoregressive dependency rather than every optimisation.
How a model generates a response Step-by-step walkthrough
Choose a step
Reduced motion is on. Choose each step manually.
1. Initialise once
The runtime loads model weights into memory before serving. They stay resident across requests while the model remains loaded.
InputNo requestResident modelWeights retainedOutputResponseModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cacheNo request cacheContext state, not learned weights. Schematic; blocks are not to scale.Illustrative inputWaiting for a request2. Send a prompt
A request brings new input, not a new copy of the weights. It can wait in a queue until the scheduler admits it.
InputWhy does memory matter?Resident modelWeights retainedOutputResponseModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cacheNo request cacheContext state, not learned weights. Schematic; blocks are not to scale.Illustrative inputWhy does memory matter?3. Prefill
The model processes the prompt. Attention keys and values for its positions form this request’s KV cache, so later tokens can use that context.
InputProcess prompt positionsResident modelWeights retainedOutputResponseModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cachePrompt contextContext state, not learned weights. Schematic; blocks are not to scale.Illustrative inputPrompt positions processed4. First token
Prefill produces the scores used to select the first output token. The response starts; model weights are unchanged.
InputPrompt consumedResident modelWeights retainedOutputNext tokenModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cachePrompt contextContext state, not learned weights. Schematic; blocks are not to scale.Illustrative outputMemory5. Decode
Feed the previous output token back through the resident model. Read the existing context, append its new keys and values, then select the next token.
InputPrevious token: MemoryResident modelWeights retainedOutputNext tokenPrevious output → resident model + cached context → next outputModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cachePrompt contextOutput 1Context state, not learned weights. Schematic; blocks are not to scale.Illustrative outputMemory matters6. Decode again
Repeat the same decode operation. Output extends one token at a time and the request’s KV cache grows as generated tokens are processed.
InputPrevious token: mattersResident modelWeights retainedOutputNext tokenPrevious output → resident model + cached context → next outputModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cachePrompt contextOutput 1Output 2Context state, not learned weights. Schematic; blocks are not to scale.Illustrative outputMemory matters.7. Finish the request
A stop condition ends generation. The runtime can reclaim this request’s cache; the model stays loaded for the next request.
InputRequest endedResident modelWeights retainedOutputResponseModel memoryWeights residentInitialisation is not repeated per request.Request memory / KV cacheNo request cacheContext state, not learned weights. Schematic; blocks are not to scale.Illustrative outputMemory matters.
Application latency is bigger than model latency.
- Coding agents
- The model returns a tool request; the harness checks authority, executes the tool, adds its result to context and calls the model again. Tool execution and dependent calls contribute to task latency. A fast token rate alone does not establish a fast or correct coding task.
Coding agents: context, tools and results Step-by-step walkthrough
Agent / tool loopStep through the architectureChoose a step
Reduced motion is on. Choose each step manually.
1. Send context
The harness sends instructions and relevant repository context to the model.
HarnessContext + permissionsContext envelopeinstructionsselected filesModel APIText + tool callsTool environmentRead / run / returnHarnessContinue or finishHarness sendsinstructions + selected filesTool result returns to the harness, then the next model call.2. Request a tool
The model returns a structured tool call. It does not run a shell or read your filesystem by itself.
HarnessContext + permissionsModel APIText + tool callsTool call envelopetool nameargumentsTool environmentRead / run / returnHarnessContinue or finishModel requestsread_file(path)Tool result returns to the harness, then the next model call.3. Execute outside the model
The harness checks permissions and runs the requested tool in its execution environment.
HarnessContext + permissionsModel APIText + tool callsTool environmentRead / run / returnTool executionpermission checkfile readHarnessContinue or finishHarness executesauthorised file readTool result returns to the harness, then the next model call.4. Return the result
The harness adds the tool result to the conversation and calls the model again. This next model call depends on that result.
HarnessContext + permissionsExpanded contextinstructionsselected filestool resultModel APIText + tool callsTool environmentRead / run / returnHarnessContinue or finishTool result joins contextfile contents → next model callTool result returns to the harness, then the next model call.5. Continue or answer
The model uses the returned context to choose another tool call or a response. The harness decides when the task is done.
HarnessContext + permissionsModel APIText + tool callsTool environmentRead / run / returnHarnessContinue or finishModel decisionnext tool callor responseModel returnsnext action or proposed answerTool result returns to the harness, then the next model call.
Agent/tool sequence. Each dependent step waits for the one before it. - Voice
- External speech recognition produces text, the text-model API generates response text, and external speech synthesis produces audio. The voice application owns turn detection, buffering and interruption cancellation. Measure first-audio latency and stale-response handling across that full chain.
Voice: speech services around a text model Step-by-step walkthrough
Conversation across three servicesStep through the architectureChoose a step
Reduced motion is on. Choose each step manually.
1. Recognise speech
A separately supplied speech-recognition service turns an utterance into text.
Speech recognitionExternal serviceAudio inputaudio chunkaudio chunkaudio chunkText modelActaserve APISpeech synthesisExternal serviceVoice applicationPlayback + turn-takingExternal speech recognitionaudio → transcriptTurn-taking and cancellation belong to the voice application.2. Generate text
Your voice application sends that transcript and conversation context to the text-model API.
Speech recognitionExternal serviceText modelActaserve APIText inputtranscriptconversation contextSpeech synthesisExternal serviceVoice applicationPlayback + turn-takingActaserve text-model roletranscript → response textTurn-taking and cancellation belong to the voice application.3. Synthesise audio
A separately supplied synthesis service turns response text into audio. First-audio latency includes the stages before it.
Speech recognitionExternal serviceText modelActaserve APISpeech synthesisExternal serviceAudio outputaudio chunkaudio chunkaudio chunkVoice applicationPlayback + turn-takingExternal speech synthesisresponse text → audioTurn-taking and cancellation belong to the voice application.4. Handle an interruption
When the user speaks again, your voice stack stops or cancels outstanding output and updates the conversation before the next turn.
Speech recognitionExternal serviceText modelActaserve APISpeech synthesisExternal serviceVoice applicationPlayback + turn-takingNew turnold output cancellednew utteranceVoice application controlscancel old response / begin new turnTurn-taking and cancellation belong to the voice application.
Conversational pipeline: separately supplied speech components around a text-model API. - Video generation
- Submit through
POST /v1/videos/generations, keep the job ID and pollGET /v1/videos/generations/:id. Handle pending, completed and failed outcomes. Polling is not a new generation request; a retry policy must avoid accidentally creating duplicate jobs.Video: submit, track and retrieve a job Step-by-step walkthrough
An asynchronous job, not a streamStep through the architectureChoose a step
Reduced motion is on. Choose each step manually.
1. Submit
Your application submits a prompt and supported generation settings. It must handle submission errors before tracking a job.
ApplicationPrompt + settingsSubmissionpromptsettingsGeneration APIJob identifierSupplier jobPending / processingApplicationResult or errorSubmit requestPOST /v1/videos/generations2. Keep the job ID
The API returns an identifier. Your application keeps it so the user can leave the submission screen while the job is pending.
ApplicationPrompt + settingsGeneration APIJob identifierJob recordjob ID retainedpendingSupplier jobPending / processingApplicationResult or errorApplication storesjob ID3. Poll status
Use the same job ID to check status while the supplier processes the job. Polling observes the job; it does not submit another generation.
ApplicationPrompt + settingsGeneration APIJob identifierSupplier jobPending / processingSame job recordjob ID retainedstatus checkApplicationResult or errorCheck statusGET /v1/videos/generations/:id4. Handle the result
When complete, use the returned result reference. A failed job needs an error state, not a result preview; pending work stays pending.
ApplicationPrompt + settingsGeneration APIJob identifierSupplier jobPending / processingApplicationResult or errorTerminal branchcompleted: result referencefailed: errorBranch on terminal statuscompleted → result / failed → error
Asynchronous video generation: submit a prompt, track the job, retrieve the result. - Robotics and sensing
- Sensor observations become perception estimates. Local placement can reduce dependence on a remote link, but safety control and actuation authority need a separate specification. A network-dependent model is not a safety interlock. This is a deployment-design discussion, not a deployed fleet service.
Robotics: perception and separate safety control Step-by-step walkthrough
Perception and control have different jobsStep through the architectureChoose a step
Reduced motion is on. Choose each step manually.
1. Capture observations
Sensors produce observations. Decide the sample rate and what can be discarded before moving a continuous stream.
SensorsObservationsObservationssensor framerange samplesLocal perceptionEstimates, not authorityRemote analysisOptional network pathSensor inputframes / range observationsSeparate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.2. Run local perception
A perception model converts observations into estimates. Keeping it near the sensor changes the dependence on network connectivity.
SensorsObservationsLocal perceptionEstimates, not authorityEstimatesobjectspositionsuncertaintyRemote analysisOptional network pathPerception outputobjects / positions / uncertaintySeparate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.3. Choose what travels
Selected observations or summaries may go to remote analysis when the link allows. Network loss must be part of the deployment design.
SensorsObservationsLocal perceptionEstimates, not authorityRemote analysisOptional network pathSelected datasummaryoptional remote analysisOptional remote workloadselected data → analysisSeparate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.4. Keep safety separate
Safety logic must independently constrain actuation. A perception estimate or remote model response is not permission to move.
SensorsObservationsLocal perceptionEstimates, not authorityRemote analysisOptional network pathSeparate control authoritysafety inputs → interlocks → actuationSeparate safety-control boundarySafety inputs → safety logic → actuator limitsNo direct actuation path from the model.Independent controlsafety inputsinterlocksactuator limits
Perception architecture. Keep safety control within a separate, independently specified boundary.
Actaserve supports coding/text/tool integration and asynchronous video generation through the current API surfaces, subject to model and supplier support. Speech services and robotics safety-control software are separately supplied.
Price the useful result, not the headline rate.
Start with a workload definition: model and licence, precision, prompt and output lengths, traffic shape, concurrency, tool use and acceptance criteria. Keep the incumbent and candidate on the same test contract. There is no matched Actaserve coding benchmark in this guide.
Keep the denominators visible.
- Time to first token
- Request start to first output. State whether the measurement includes client networking and queueing, and report the distribution rather than only its average.
- Output tokens per second
- Generation rate after output begins. Label per-request versus aggregate throughput. Neither describes prompt processing time or tool delays.
- Accepted tasks per period
- Completed tasks that pass the agreed quality and safety criteria. Include retries and failures in the cost numerator even when they do not count as accepted output.
Cost per accepted task = total attributable cost ÷ accepted tasksEnergy per accepted task = measured joules ÷ accepted tasksIf no tasks pass, cost per accepted task is undefined—not zero.
Compare complete operating arrangements.
For an API, account for billed usage, retries and tool or storage services outside the model charge. For dedicated capacity, include the reservation or hardware cost over its useful life, finance where relevant, energy, facility costs, network, software, staffing, support, spares and downtime. Avoid counting the same cost in both a bundled service fee and a separate line item.
Utilisation spreads fixed costs over useful work. At a fixed period cost, halving accepted task volume doubles fixed cost per accepted task. That arithmetic is not a performance forecast: raising utilisation can also increase queues and miss latency targets. Leave headroom for bursts, failures and maintenance.
Report demand assumptions, idle periods, acceptance rate, depreciation period, residual value and what the meter excludes. Evaluate a range of demand levels instead of treating a single full-utilisation estimate as the business case.
Run acceptance before committing capacity.
- Pin the checkpoint, runtime version, precision and sampling settings.
- Use representative task inputs with authorised data and explicit pass criteria.
- Record warm and cold conditions separately, including concurrency and request-length distributions.
- Measure first response, complete-task latency, quality, failures, memory and power over a matched window.
- Repeat under expected sustained traffic and bursts. Record what happens during supplier, network and host failures.
Only measured results from that configuration can establish whether it meets your workload’s targets. Manufacturer specifications establish equipment characteristics, not service outcomes.
Put sovereignty in the operating terms.
A hardware brand does not establish data residency, isolation or ownership. Specify processing, storage, backups and permitted transfers; administrator and root access; support access; tenant isolation; maintenance and incident response; and the model and data handover process.
Dedicated compute reservations are open. Delivery is scoped per project: Actaserve will own and operate the hardware, including procurement, colocation power/cooling, model bring-up and ongoing operations. Customers will choose root/bare-metal access or a managed API endpoint. Private-weight deployments will be custom quoted, subject to model, licence and runtime qualification.
API access does not reserve a particular chip. Dedicated hardware availability, location, capacity, price, lead time and access terms are agreed per project. The architectures here are illustrative, not an availability guarantee.
Bring the workload definition and the boundaries you need to protect. The configuration discussion should end in explicit delivery and acceptance terms, not an assumed hardware guarantee.
Discuss your compute requirementsSources and reading notes.
This guide explains general architecture mechanisms and measurement boundaries, not a particular product’s specifications or measured performance. Hardware examples, manufacturer figures and manufacturer citations are not part of this public guide.
The step-driven explanation approach was informed by MTS Compute. Its models, calculators and performance data are not Actaserve evidence. Here, weight loading is explicitly separated from per-request prefill and decode.