AI Provisioning & Setup
Every eRacks AI server already ships AI-ready, free

The basic preinstall is included with every eRacks AI system, always has been, always free: your choice of open-source OS plus the local-AI stack - Ollama, Open WebUI (the familiar chat interface), vLLM, llama.cpp, PyTorch and the Hugging Face toolchain - installed, updated, and tested before it leaves our facility. Power it on, load a model, go.
AI Provisioning & Setup: from AI-ready to AI-in-production, $1,795

The paid package is everything between a working stack and a working deployment, tailored to your vertical and your environment:
- Model strategy for your use case: we select and install the right open models for your work - legal, healthcare, engineering, support, research - from the current best of DeepSeek, Llama, Qwen, Mistral and friends, including MoE (mixture-of-experts, the architecture behind DeepSeek-class models) where your hardware supports it.
- Quantization & VRAM fit: models quantized (GGUF, AWQ, FP8 - compressed to fit more capability into your GPU memory) and verified against your actual VRAM budget, not a spec sheet.
- Serving tuned to your hardware: vLLM/llama.cpp configs set for real throughput - context length, continuous batching, tensor parallelism across your GPUs - with benchmark numbers (tokens/second, latency) recorded and handed to you.
- Private by design: air-gapped or egress-controlled network configuration, so your data and prompts never leave the building.
- Your documents, searchable: a starter RAG setup (retrieval-augmented generation - the model answers from YOUR files) wired into Open WebUI.
- Beyond text: image generation (ComfyUI with Stable Diffusion/FLUX-class models) and speech-to-text (Whisper) configured on request.
- Multi-node AI clusters: when one server is not enough, we design and deploy clustered inference across several machines - Ray (the distributed compute framework that ties the nodes into one), vLLM distributed serving and llama.cpp RPC for tensor and pipeline parallelism - so a single very large model, or many simultaneous users, spreads across the whole cluster. It is a discipline we practice on our own hardware, not just a slide.
- Handoff, not lock-in: a live two-hour working session with your team, a written record of every config, and 30 days of follow-up email questions included.

Pair it with any eRacks AI system, or run it remotely on capable hardware you already own. Need a multi-node AI cluster, fine-tuning pipelines, or SSO integration? Select the Extended tier above and we quote it to your exact setup. Ongoing assistance is always available after the sale as a consulting add-on, no lock-in, like everything we do.
Configure AI Provisioning & Setup
Choose the desired options and click "Add to Cart", or "Get a Quote". Please add any additional requests and information in the "Notes" field. Your quote request will be sent to your profile's eMail if you are logged in, otherwise enter the email address below (required only if not logged in).
Current Configuration
Default Configuration
eRacks Open Source Systems