Private AI · managed operation

Private AI infrastructure and digital workers.

We design, deploy, and operate AI environments that can understand digital work, use approved tools, and continue scheduled jobs. Start with one measured workload, then choose local, managed-cloud, hybrid, or air-gapped operation.

Workload
Policy boundary

Governed runtime

See · act · verify

Approved tools
Outcome check
  • One measured workload first
  • Human control for sensitive actions
  • Capacity proven before scale
  • Local and approved-cloud options

Operating boundary

Deployment architecture.

The same workflow can use different infrastructure. The right choice follows the data, latency, integration, support, and control requirements—not a preferred vendor.

Local

01

Model serving and selected data stay on equipment in your environment. Egress, updates, and operator access are configured for that deployment.

Best when
The workload uses sensitive internal data or needs predictable on-site operation.
Plan for
Hardware ownership, capacity, patching, physical access, and recovery.

Managed cloud

02

Dedicated or shared cloud resources with provider, region, retention, and administrator access agreed before launch.

Best when
Elastic capacity and managed operations matter more than equipment ownership.
Plan for
Provider terms, regional availability, cost controls, and service dependencies.

Hybrid

03

Keep selected retrieval and tools private while routing approved tasks to chosen external model or voice services.

Best when
A team wants private company context with access to selected cloud experiences.
Plan for
Data minimisation, routing policy, failure modes, and explicit external-call boundaries.

Air-gapped isolation

04

Isolated operation is available for compatible workloads. Model delivery, updates, integrations, and support follow a separate offline process.

Best when
Network isolation is a hard operating requirement and the workflow fits offline execution.
Plan for
Offline supply chain, update windows, restricted integrations, and on-site support.

Digital work loop

Digital worker operating model.

Each worker executes a repeatable task in a bounded environment with inspectable decisions, controlled actions, and a defined handover to a person.

  1. 01

    See

    Interpret approved screens, documents, images, and application state.

  2. 02

    Choose

    Apply task policy, context, and confidence thresholds before selecting a next step.

  3. 03

    Act

    Use computer control, APIs, MCP servers, and business systems through scoped identities.

  4. 04

    Verify

    Check the resulting state, record tool events, and request confirmation where policy requires it.

  5. 05

    Continue

    Run schedules and durable jobs with bounded retries, alerts, and a defined human handover.

Computer control

A worker can inspect an approved desktop or browser session and operate software through a constrained environment when direct APIs are not suitable.

Control: High-impact actions can require confirmation; sessions, applications, and credentials remain explicitly scoped.

Tools and company data

Connect selected APIs, MCP servers, databases, document stores, and business systems to the task—not to an unrestricted general agent.

Control: Each connector receives the minimum identity, data scope, and tool permissions needed for its job.

Scheduled operations

Run recurring research, reporting, monitoring, reconciliation, preparation, or support work with observable job state.

Control: Retries, alerts, escalation, stop conditions, and ownership are designed before unattended operation.

Voice and interface

Voice and conversation interfaces.

Interface choice and data boundary are separate architecture decisions. Your team can use the experience that fits the job without giving that interface unrestricted access to private tools or data.

01 / LOCAL

Local voice

Use a suitable local speech stack when measured latency, language quality, hardware capacity, and support requirements fit the workflow.

Candidate models and voices are demonstrated and accepted for the client's language, environment, and device before commitment.

02 / APPROVED CLOUD

Existing cloud experiences

A client-approved service such as ChatGPT can remain the conversation layer while the private worker exposes only authenticated, policy-checked actions.

Provider retention, regional processing, permissions, and transmitted fields are reviewed separately from the private runtime.

System design

Control plane and system boundaries.

Each layer has an owner, an access boundary, and evidence. The exact products are selected after the workload and operating constraints are understood.

  1. Entry

    Web, chat, voice, and schedules

    The interfaces through which people and systems start approved work.

  2. Policy

    Identity, permissions, approvals, and tool allow-lists

    The rules that decide what may happen, with which identity, and when a person must intervene.

  3. Runtime

    Selected models and isolated work sessions

    Local or approved cloud inference combined with the execution environment required by the task.

  4. Systems

    Explicit APIs, MCP connectors, and data scopes

    Only the company systems and data paths needed for the measured workload.

  5. Evidence

    Outcome checks, logs, alerts, and handover

    Signals that show what the worker attempted, what changed, and who owns an exception.

Privacy and data-retention controls.

A reference architecture is not a certification.

  • 01A zero-retention requirement is checked across each provider, log, support path, and backup policy; it is not assumed for every deployment.
  • 02EU residency can be selected where the chosen provider and service support it, then documented in the deployment boundary.
  • 03Data minimisation or anonymisation can be placed before approved external calls where the workload allows.
  • 04Sensitive actions can require human confirmation, and every integration receives only the scope it needs.

Workload sizing

Hardware selection.

Model memory is only the first constraint. We measure quality, prompt and context size, output rate, concurrent activity, latency, runtime compatibility, power, support, and recovery before recommending a system.

01

Compact Local AI

32 GB unified

AMD Ryzen AI Max PRO 380-class workstations.

Candidate fit: Private retrieval, speech, screen understanding, and narrow multimodal workflows after task-level testing.

Dated evidence for this tier only

32 GB retail reference

02

Team Local AI

64 GB unified

AMD Ryzen AI Max+ 395-class compact systems.

Candidate fit: Quantized 27B-class candidates such as Qwen3.8-27B when the selected runtime, quantization, context, and measured latency fit the workload. AMD documents Gemma 3 27B QAT on this memory tier.

Dated evidence for this tier only

64 GB retail reference

03

AI Studio

128 GB unified

AMD Ryzen AI Max+ PRO 395-class workstations and compact systems.

Candidate fit: Larger low-bit candidates, longer contexts, multimodal evaluation, and several local services when a workload benchmark supports the configuration.

Dated evidence for this tier only

128 GB retail reference

04

NVIDIA DGX Spark

128 GB unified

GB10 CUDA desktop development and inference systems.

Candidate fit: NVIDIA publishes support for testing or inference up to 200B parameters and fine-tuning up to 70B. These are capacity ceilings, not latency, context, concurrency, or SLA promises.

Dated evidence for this tier only

DGX Spark retail reference
NVIDIA DGX Spark compact AI system

05

NVIDIA DGX Station

748 GB coherent

GB300 Grace Blackwell Ultra deskside systems, configured and sourced by quotation.

Candidate fit: NVIDIA publishes support for models up to one trillion parameters and up to seven isolated MIG instances. The final checkpoint, precision, context, throughput, site, and support requirements remain quote-time decisions.

Dated evidence for this tier only

Station-class retail reference

Capacity, honestly stated

Worker slots are not employee equivalents.

We size simultaneous responsive work and scheduled duty cycles, then map those measured slots to named workflows. We do not turn a token-rate estimate into a claim about human headcount.

concurrent responsive sessions ≈ measured task throughput ÷ required throughput per active session

Example only: if the accepted workload sustains 40 output tokens per second and one interactive session needs 10, that is roughly four simultaneous sessions. At a 20% active-duty assumption it may support about twenty assigned routine workflows—but still not twenty human-equivalent employees.

Illustrative platform families reviewed 22 August 2026. Availability and specifications change. Every model/runtime combination is benchmarked before sizing.

Evidence boundary

Hardware is selected per workload.

The links inside the five tiers are dated examples that substantiate the class of system being discussed. They are not a public hardware catalogue, live inventory, or Algovectra offer.

Current price, availability, warranty, delivery, configuration, security hardening, and managed service are confirmed only in a project quotation.

Platform scope

Integrations.

Identity, tools, orchestration, and observability are selected components—not an automatic bundle of platforms or licences.

Identity and access
OIDC, SAML, service identities, conditional access, and role-based policy where required.
Tools and secrets
MCP or API connectors, secret brokering, allow-lists, sandboxing, and confirmation rules.
Workflow runtime
Durable jobs, queues, schedules, bounded retries, and human-in-the-loop states using the selected platform.
Observability and handover
Task outcomes, model and tool events, alerts, runbooks, ownership, and agreed retention.

Delivery sequence

Pilot scope.

Timing is confirmed after discovery because hardware availability, security review, integrations, and acceptance scope change the work.

  1. 01

    Discover

    Define the task, current system, data boundary, human roles, and acceptance test.

  2. 02

    Prove

    Run the smallest representative workflow and measure quality, latency, failure modes, and cost.

  3. 03

    Integrate

    Add only the identities, tools, data paths, monitoring, and controls needed for the accepted workflow.

  4. 04

    Operate or hand over

    Document ownership, service targets, updates, recovery, and the agreed managed or client-operated model.

Contact

Discuss a Private AI workload.

Define the task, data boundary, tools, human controls, and acceptance test before choosing hardware or models.

Book a scoping call