Tolka Edge SymbolTolka Edge WordmarkDocs

Rig Lifecycle

The full lifecycle of a dedicated AI Rig — booting, active serving, and termination.

Every dedicated AI Rig moves through a well-defined lifecycle. Understanding it helps you predict cold-start latency and manage your wallet balance.

Lifecycle at a glance

Provisioning
allocate GPU
Loading
weights to VRAM
Active
serving
Terminated
offline

1. Provisioning

When you deploy a new AI Rig from the dashboard, Tolka Edge acquires a lock and provisions a dedicated GPU instance on our cloud infrastructure. Billing for the Rig starts immediately at this stage, as the GPU is allocated to your deployment. We maintain warm container images for Qwen3-14B, so you don't have to wait for massive Docker pulls.

2. Model loading

Once the instance boots, the vLLM engine initializes and loads the Qwen3-14B weights into the GPU VRAM. Once the model responds to an internal health probe, the node transitions to Active.

3. Active & Serving

While active, your dedicated Rig serves your requests with low latency. Tolka Edge continues to bill your wallet for the active GPU time, measured to the second.

4. Termination & Auto-shutdown

Because your Rig is a dedicated instance, it remains running until one of the following occurs:

Manual Termination

You explicitly terminate the Rig from the Dedicated GPU Dashboard when you are done.

Wallet Depletion (Auto-shutdown)

If your wallet balance runs out, Tolka Edge will automatically trigger a shutdown of your Rig to prevent negative billing.

No automatic pause

Tolka Edge does not automatically pause or resume your Rig based on idle time. You are billed for the entire time the Rig is in the Active state. Be sure to terminate your Rig when you no longer need it.