Skip to main content
On-premise AI, delivered as one project

Your own AI agents, inside your own walls.

For companies that cannot send documents, code or customer data to a public AI service. We size the hardware, install and tune the models, put the agent software on every seat, and stay for support.

What you get

One project, four deliverables.

  1. 1

    An inference server on your GPU

    vLLM or Ollama on an RTX 4090, RTX 5090 or DGX Spark, or on GPUs you already own. Tuned for your team size and benchmarked with your own prompts before go-live.

  2. 2

    Nucleus Work on every desk

    The desktop agent for business users, pointed at the office model server. Data stays in one local database file per machine; no accounts to manage.

  3. 3

    The NE CLI for developers

    A terminal coding agent for every engineer, with the four-agent team mode, skills and MCP, using the same model server.

  4. 4

    Rollout and support

    Onboarding sessions, a shared MCP and skills setup, model updates, software upgrades and a support line.

Hardware

Three cards cover most teams.

Memory decides which model fits; concurrency decides how many people it serves well. The table is our starting point, and the sizing step replaces it with your numbers.

GPUMemoryWhat fitsTeamBasis
NVIDIA RTX 409024 GBModels up to about 14B parameters at full speed; larger ones quantizedA small team doing documents, email and chatGuidance
NVIDIA RTX 509032 GBQwen 3.8 27B in NVFP4 on vLLM, long sessions with a CPU cache tier8 developers per card for interactive coding agents; 16 workableMeasured
NVIDIA DGX Spark128 GB unified70B-class models and long context on one desk-side boxLarger teams, or bigger models for fewer peopleGuidance

Measured with Nucleus agents on one RTX 5090: 8 developers at about 57 tokens per second each with all sessions resident; 16 developers all finish their tasks at roughly 1.5 times the time.

Nucleus benchmark, 2026-09-01/02, Qwen 3.8 27B NVFP4 on vLLM 0.28

We size per team from your prompts, context lengths and headcount. Existing GPUs and multi-GPU servers are welcome.

Offline or hybrid

Decide what may leave the network.

Fully offline

For regulated or confidential work. Public egress can be blocked at the firewall.

Stays inside

  • Prompts, files and conversations never leave the LAN
  • Model weights live on your server
  • Sessions and settings stay in local files on each machine
  • Updates are delivered as files you apply

Leaves the network

  • Nothing

Hybrid with Nucleus AI Cloud

Local by default, with hosted models on Nucleus AI Cloud for approved people and workspaces, through the Nucleus gateway.

Stays inside

  • Sensitive workspaces stay on the local model
  • Files and sessions stay local either way
  • Provider credentials stay with the admin

Leaves the network

  • Prompts and context for workspaces that choose a gateway model, within per-member quotas
Security posture

What the software does, stated plainly.

These are the properties of the shipped software. Your security team can verify each one on a test machine before rollout.

One local database file

Nucleus Work stores users, sessions and messages in a single SQLite file on the machine. No Postgres, Redis or cloud database.

No login, no cloud account

The desktop app starts straight into the workspace. Cloud linking exists but stays off unless a user links an account.

Sandboxed plugins

Every agent module runs as a WebAssembly plugin inside the host, talking through a narrow, JSON-only interface.

Keys stored privately

Model keys are written with private file permissions and referenced by path. They are never placed in a config file.

Loopback-only previews

Apps that Nucleus Work builds preview on 127.0.0.1 of the same machine. Nothing is exposed to the LAN or the internet.

How an engagement works

From first call to a working team.

Most deployments take weeks, not quarters. You keep the hardware, the models and the data; we do the sizing, the installation and the tuning.

  1. 1

    Assess

    A short workshop on the workflows you want automated, the data involved, your network policy and how many people will use it.

  2. 2

    Size and quote

    We propose the GPU, the model and the deployment shape for your headcount. Hardware you already own is fine.

  3. 3

    Install and tune

    We set up the inference server, load the models, benchmark with your prompts, and put Nucleus Work and the CLI on every seat.

  4. 4

    Roll out and support

    Team onboarding, model updates, software upgrades and a support line for as long as you want one.

Questions

Before you write to us.

Do we need an internet connection?

Not for day-to-day use. Fully offline installs run with public egress blocked. Internet is only needed if you choose the hybrid option, and then only for the workspaces that use a gateway model.

Can we still use hosted models?

Yes, in hybrid mode. An admin connects the subscription your company owns to Nucleus AI Cloud, sets per-member quotas, and the hosted models appear in the picker in Nucleus Work and the CLI.

Which local models do you install?

Open-weight models served with vLLM or Ollama. Our current reference is Qwen 3.8 27B on an RTX 5090. We choose per team from your prompts, languages and context length.

Can you use hardware we already own?

Yes. We assess it in the sizing step. Anything with an NVIDIA GPU and enough memory for the chosen model is a candidate, and multi-GPU servers work with vLLM.

Do you also run the web platform and sandboxed agents on our servers?

That is a scoped project rather than an off-the-shelf install. If you need the hosted platform experience inside your network, we discuss it at the assessment and quote it separately.

Where are you based?

Singapore. We work on site in Singapore and remotely across the region.

Tell us about your team.

A few lines about what you want automated, roughly how many people would use it, and whether you already have GPUs. We reply with a proposed shape and next steps.

Do you already have GPUs?

Or email us directly hello@nucleusenterprise.ai