Aug 27, 2026 - eRacks Ships Benchmarked Dual Intel Arc Pro B70 AI Servers and Prices AI Provisioning at $1,495

Campbell, CA - Aug 27, 2026

FOR IMMEDIATE RELEASE

eRacks Ships Benchmarked Dual Intel Arc Pro B70 AI Servers and Prices AI Provisioning at $1,495

The open-source server maker published tokens-per-second numbers measured on its own bench - 54 tokens per second on Qwen3-14B, 26 on Qwen3.6-27B from a single card - and made the provisioning behind them a standard priced offering: $1,495, $2,495 with a private RAG stack, included on flagship orders.

eRacks Open Source Systems today published benchmark results measured on a dual Intel Arc Pro B70 server during its pre-ship provisioning pass, and announced that the provisioning work behind those numbers is now a standard priced offering across its AI server line.

The measured numbers, from the company's bench this week: Qwen3-14B generating 54 tokens per second, and the larger Qwen3.6-27B holding 26 tokens per second on a single B70, steady across back-to-back 300-token responses. A token is roughly three-quarters of a word, so the larger model writes faster than most people read. The serving stack is fully open source: llama.cpp's official Intel build running in rootless Podman containers (no root daemon), exposing the industry-standard OpenAI-compatible API, with models resident entirely in GPU memory.

The hardware is the point. The Arc Pro B70 carries 32GB of VRAM (the GPU's onboard memory, the hard limit on what models fit) per card, so a two-card server fields 64GB of GPU memory for less than the list price of a single 96GB flagship datacenter card - which the company's purchasing records, published August 21, put at $16,000 and climbing. For private AI (running models on your own hardware, on your own data), that ratio of memory to dollars is the value play of 2026.

"Benchmarks on a spec sheet are marketing. Benchmarks on your machine are engineering," said Joseph Wolff, founder and CTO of eRacks Systems. "Getting these cards to production took three fixes you will not find in any manual - GPU power management that puts cards to sleep permanently, container networking that resets every connection while the server looks healthy. We solved them on the bench, and every AI server we ship now leaves with its own measured numbers and the rebuild notes in the customer's hands."

That work is now a named product: eRacks AI Provisioning & Setup covers burn-in, GPU bring-up with every fix applied, deployment and benchmarking of the customer's chosen models on the customer's actual hardware, and full rebuild documentation - $1,495, or $2,495 including a private RAG stack (retrieval-augmented generation: a chat interface plus vector database that lets the models answer from the customer's own documents, entirely offline). It is included at no charge on flagship orders.

The line ships configured to order at live prices: eRacks/AIDAN with one B70 from $13,895, eRacks/AINSLEY with two B70s and 64GB of GPU memory - the configuration class benchmarked above - from $21,395, the full AI server line from $7,695, and the 8-GPU eRacks/HIGHLANDER flagship from $154,995. Buyers can compare ownership against their current cloud or subscription spending with the company's no-signup calculator at https://eracks.com/tco/ and configure any machine at https://eracks.com/products/ai-rackmount-servers/

About eRacks Open Source Systems

eRacks Open Source Systems, founded in 1999 and headquartered in Campbell, California, specializes in open-source servers and storage. The company builds custom-configured rackmount servers, NAS (network-attached storage) systems, HPC (high-performance computing) clusters, and AI inference servers, all shipped with Linux or other open-source operating systems. More at https://eracks.com

Media Contact

Joseph Wolff eRacks Open Source Systems joe@eracks.com https://eracks.com