Introduction
Tolka Edge is a one-click platform for deploying dedicated GPU AI Rigs with production-ready, OpenAI-compatible inference APIs.
Tolka Edge is a one-click platform for deploying dedicated GPU AI Rigs with production-ready, OpenAI-compatible inference APIs.
Tolka Edge provides One-Click AI Rigs. Deploy a dedicated GPU running Qwen3-14B through an optimized vLLM runtime, then use it through a single OpenAI-compatible API endpoint.
Tolka handles GPU provisioning, model deployment, runtime configuration, lifecycle management, and usage billing so you can run dedicated inference without managing GPU infrastructure yourself.
⚡
Quick Start
Deploy a Rig and make a request.
🔑
Authentication
Create keys and authorize requests.
📘
API Reference
Endpoints, parameters, and error codes.
What is Tolka Edge?
Most teams that self-host open models have to manage GPU provisioning, CUDA/runtime compatibility, model serving, scaling, and billing infrastructure. Tolka Edge abstracts these operational details behind a simple deployment and API experience.
- One-Click Deployment. Select a supported GPU and model, deploy your Rig, and Tolka provisions the infrastructure and starts the inference server automatically.
- Dedicated Compute. Your model runs on a dedicated GPU instance rather than shared inference capacity.
- Transparent Billing. Rigs are billed continuously based on actual active GPU time, with usage tracked to the second.
Why Tolka Edge?
⚡
High-Throughput Inference
Qwen3-14B is served with AWQ quantization and optimized vLLM inference. In internal testing, the deployment reached approximately 82 tok/s single-stream decode throughput on an RTX 3090.
🧩
OpenAI Compatible
Use the OpenAI SDK or existing frameworks such as LangChain and Vercel AI SDK by changing the base URL and API key.
🖥️
Managed Lifecycle
Tolka handles Rig startup and termination, with optional automatic shutdown timers to help control GPU costs.
📊
Transparent Billing
Pay for actual active GPU time with usage tracked to the second.