Skip to main content

[ 00 / 08 ]

[ PRIVATE ASSISTANTS | YOUR DOCUMENTS | YOUR INFRASTRUCTURE ]

A PRACTICAL START. CONTROL OVER YOUR DATA. ROOM TO GROW.

SOVEREIGN AI | BUILT FOR YOUR BUSINESS

Put AI to work. Keep control of your data.

Find answers in company documents, prepare drafts and put AI to work with your business data, on your own infrastructure.

Built for small businesses and growing teams. We bring the hardware, models and integration together around your needs and budget.

START WITH THE WORK YOU WANT TO IMPROVE

Less searching. Less repetitive work. More useful AI.

Sovereign AI means choosing where your data is processed, who can access it and which systems you depend on. The starting point is a useful business application, not a hardware specification.

Find answers in your company knowledge

Let your team search manuals, procedures and project documents through a private assistant. We scope the data connections, permissions and source references around the information people actually need.

Give everyday work an AI assistant

Summarise documents, prepare drafts and support customer follow-up. Connect a private AI environment to your workflows, with AIDesk as one application option.

Explore the AIDesk workspace

Build and code with local AI

Give developers an internal AI endpoint for coding assistance and business tools. Choose a model for your tasks and validate it with your code, documents and expected users.

RIGHT-SIZED, NOT RACK-SCALE BY DEFAULT

Start with what your team needs, not a cluster.

A private assistant, document search or local coding endpoint has different needs from large-scale training. We size memory, compute, storage and concurrent users together, then check the result against your actual tasks.

A local AI workstation

A focused starting point for an individual or small team. Explore local assistants, coding and content workflows without buying a cluster.

A shared on-prem AI server

Serve internal applications through a local API. Size GPU memory and request capacity for the models, context lengths and response times your team needs.

A measured expansion path

Add capacity when utilisation and demand justify it. If your workload needs high-density, multi-GPU infrastructure, we have a separate cluster offering.

Explore GPU servers & clusters

ONE EXAMPLE, NOT A CLUSTER-SIZED COMMITMENT

You may need one workstation, not a data center.

A Radeon AI PRO R9700 with 32 GB of GPU memory is one possible starting point for a local AI workstation. We select the complete system around the model, the people using it and the response times you need. A suitable system matters more than the biggest GPU.

AI-generated concept illustration, not a manufacturer photograph or an exact supplied configuration. System configuration and availability are confirmed per project; the calculator uses a complete-system allowance, not the price of this GPU.

MORE THAN A HARDWARE DELIVERY

A working solution, not a box of parts.

We scope a complete deployment: suitable hardware, a configured model and runtime, application connections and an agreed operating plan. You know what is included and what must be demonstrated before rollout.

01 / Prove the use case

Start with a task, representative examples and the people who will use the system. Agree what a useful answer looks like, how fast it must arrive and which data may be processed.

02 / Fit the model to the system

Select the model and hardware together. Where useful, quantization reduces model memory use; fine-tuning adapts behaviour to a specific task. We check the result instead of assuming a smaller model is good enough.

03 / Integrate and hand over

Connect your applications, test real requests and simultaneous users, and document the setup. Agree access, updates, monitoring, backup responsibilities and support before it becomes a daily business tool.

BUILT ON HANDS-ON WORK

Local inference, with public work to back it up.

Our public paiton-vllm-plugin brings Paiton execution into vLLM. Its current Radeon focus is RDNA 4 on the Radeon AI PRO R9700. Support is specific to the documented model and hardware profiles, not a promise that every model runs on every GPU.

paiton-vllm-plugin

Inspect the public integration, supported profiles and setup instructions. The Paiton compiler remains proprietary; public components have their own licences.

View the GitHub repository

GGUF inside vLLM, on one Radeon

A qualified local text and image endpoint using the original GGUF weights. Read the measured results and deployment limits.

Read the local inference benchmark

Paiton performance work

Explore workload-specific results across language, image and video generation. We use measurements to guide system choices, not blanket speed claims.

Paiton

CLOUD TOKENS VS LOCAL OWNERSHIP

What could local AI save you?

Enter your monthly token volume. Compare a cloud API bill with the modeled cost of one local system, including hardware, electricity and ongoing operations.

All amounts in USD. Editable planning assumptions, not a hardware quote or guaranteed savings.

This is an estimate and can vary based on usage, the model and the applicable rates. Adjust the token volumes and prices to match your situation.

Power & capacity assumptions

Example rates, not R9700 benchmark claims. Replace them with measured rates for your model, context and workload. A 30-day month is powered on 24/7; unused hours consume idle power. The 500-hour default leaves headroom for peaks and maintenance.

A cost scenario, not a like-for-like model benchmark. The cloud service and local system use different models. A local model must first meet your quality, context, security and concurrency requirements. Token counts can differ by tokenizer.

How this is calculated

Cloud = input millions × input rate + output millions × output rate. Local = system and setup cost ÷ ownership months + active/idle electricity + monthly upkeep. Active hours = input tokens ÷ input speed + output tokens ÷ output speed, converted to hours. Cash payback = upfront cost ÷ (cloud bill minus local electricity and upkeep).

Standard uncached text pricing; include billed reasoning in output tokens. No batch discounts, cache pricing, taxes, financing, residual value, cloud tools or network fees. Add deployment, evaluation and fine-tuning costs to setup, and cooling or ongoing labour to upkeep where applicable. All-local substitution and steady monthly demand are assumed; burst latency is not guaranteed.

SOVEREIGNTY IS A DESIGN REQUIREMENT

Keep control where it matters.

Data and access

Decide where prompts, documents, outputs and logs live, who can access them and which network connections are permitted. Local inference can keep requests on your own infrastructure.

Models and dependencies

Review model licences, runtime dependencies, telemetry and update paths. Downloading models and installing software may still require external access; an offline deployment needs explicit preparation and testing.

Operations and responsibility

Agree who handles security, backups, patches and recovery. On-prem hardware alone does not guarantee sovereignty, security or regulatory compliance. The operating model has to support your requirements too.

HOW YOUR PRIVATE AI FITS TOGETHER

Inside your infrastructure
  1. Your applications

    Documents · Coding · Internal tools

  2. Your private AI

    Selected model · Secure access · Local processing

  3. Your hardware

    Workstation or on-prem server

Illustrative deployment. Access, logging and external dependencies are defined for your project.

A CLEAR DECISION, NOT AN AI LEAP OF FAITH

Is local AI the right choice for you?

Is local AI always cheaper than cloud AI?

No. Regular use and clear data-control requirements can make a local system attractive. Occasional use or highly variable demand may favour cloud services. Compare hardware, electricity, support and model quality, not just the token price. The calculator is a planning estimate, not a savings guarantee.

Do we need our own AI specialists?

You do not need to choose the entire stack yourself. ElioVP can scope the hardware, models and integration with you. Your organisation still needs an owner for the application and its data, with support and operational responsibilities agreed for the project.

Can the system run without sending prompts to a cloud AI service?

Yes, in a deployment configured for local processing. We review document extraction, model processing, search indexes, logs and external connections. Fully offline operation requires additional preparation and testing; installing a GPU alone does not create that boundary.

Will it do everything a large cloud model does?

Not necessarily. The right model depends on language, reasoning, document length, tools and available memory. We validate the tasks that matter to your business and make the trade-offs explicit before recommending a system.

Tell us what you want to run locally.

Tell us which task takes too much time, how many people need access and what must stay private. If you have a budget range, include it. We will help define a practical starting point and what to test first.