Skip to main content
Token Farm · Model Serving · AI Products

Helix

In-house inference clusters and KV cache infrastructure —
for faster, cheaper, and steadier long-context token supply.

helix · multi-agent workspace
Helix Code multi-agent workspace running parallel tasks

Token Farm: self-hosted inference infrastructure

From GPU clusters to disk-level KV cache — we rebuilt the cost structure of long-context inference, so every token costs less to produce.

🖥️

In-House GPU Cluster

A100-class clusters already running in production; dual vLLM / SGLang stacks, with long-context requests as everyday workloads.

💾

Cascade Disk-Level KV Cache

Lifts KV cache from GPU memory to NVMe and cluster storage; cache reuse skips repeated prefill and cuts memory footprint by 50–80% (design target).

🔀

Inference Gateway & Scheduling

A unified multi-model gateway: request routing, quota billing, and key & user management — turning cluster capacity into deliverable, metered token supply.

Request intakeGateway routing & quotaKV cache match (hit = compute saved)GPU inference (vLLM / SGLang)Streaming output
20.8×
TTFT Speedup
Public benchmark (cache hit)
96%
Cache Hit Rate
Repeated-prefix reuse
50–80%
GPU Memory Saved
Design-target range
20 GB/s
KV Read/Write Bandwidth
Backed by TB-class cluster throughput; depends on network

* Speedup and hit rate are measured on public benchmarks; KV bandwidth depends on network, backed by TB-class aggregate storage throughput.

Explore the platform & infrastructure →

We build and run models

From deployment and tuning to production — we operate a full model stack on our own clusters.

🚀

Running Models in Production

DeepSeek V4 Flash and other large models run in production on in-house clusters, with 400K-class server-side context configuration — long context as an everyday workload.

🛠️

Full-Stack Inference Tuning

Dual vLLM / SGLang stacks; KV sharding, quantized deployment, and tiered storage (POSIX / GDS-NvFile) tuned layer by layer for higher memory efficiency and throughput.

🔌

Multi-Model Access & Switching

A unified model abstraction and policy layer: multi-vendor model access, gradual switching, and quota control behind a single API.

Inference cost is deciding who wins in AI

Tokens are becoming AI’s most basic commodity — and their price is set by infrastructure.

📈

Inference demand is going exponential

Agents, long documents, retrieval, and multi-turn collaboration dominate real workloads — context length and concurrency are exploding together, and inference compute is the real bottleneck.

💸

Cost is the deciding factor

Model quality is converging; the bottleneck shifts from whether it works to cost per token — and GPU memory is the most expensive link in the chain.

Cache reuse is the structural answer

Every cache hit is compute saved. Make KV cache shareable, schedulable, and reusable across nodes — and token costs drop structurally.

Two agent product lines, built on our own infrastructure

AI products powered by our inference stack: one for developers, one for enterprises.

Developer · Helix Code

🧩Helix Agent

Parallel, observable, long-running AI engineering workflows: three-layer multi-agent orchestration, concurrent SubAgents, and cache + compact context management that keeps long tasks moving.

Parallel agentsMulti-remote workspacesDesktop / Mobile / Web
Explore Helix Code
Enterprise · Helix Flow

🏢Helix Flow

AI-native workplace collaboration plus an enterprise agent platform: messaging, tasks, approvals, and an AI assistant woven into daily work — with private deployment, secure connectivity, and centralized governance.

AI-native workplacePrivate deploymentSecurity & governance
Explore Helix Flow

What Makes Helix Different

Built for teams actually running engineering workflows — not one-off chat completions.

🧠

Three-Layer Multi-Agent Architecture

A manager agent guards user intent, execution agents drive the work, and SubAgents dive into sub-tasks.

Concurrent Agent Scheduling

Tool calls and SubAgent tasks run in parallel — multi-module investigations finish markedly faster than serial chat loops.

👁️

Visible Parallel SubAgents

While the main conversation keeps moving, every sub-task reports its state and results. Parallelism is no longer hidden.

📦

Cache + Compact Context Management

Large tool outputs are cached and retrieved on demand, long histories compact automatically — keeping long sessions stable.

🌐

Multi-Remote Workspaces

Local projects, remote hosts, and cloud containers under one roof — workspace-level isolation with a consistent workflow.

🔧

LSP + Tools + Skills, Unified

Code intelligence, terminal, Git, MCP tools, and Skills all flow through one execution loop, from analysis to delivery.

See Helix in Action

Real screenshots from the desktop and mobile apps.

Curious about the Token Farm, model serving, or partnerships?

Inference infrastructure, model deployment, or agent products — get in touch.