Worker types
What you’re billed for
Your total cost includes compute time and storage:Compute cost breakdown
Workers incur charges during three phases:- Start time: From when the worker starts running until it is ready, mostly your handler loading models into GPU memory. Pulling the container image and downloading cached models happen before the worker starts running and are not billed. Minimize with FlashBoot or model caching.
- Execution time: Processing requests. Set execution timeouts to prevent runaway jobs.
- Idle timeout duration: The time a worker remains active (running) after completing a request, waiting for additional requests before scaling down (default: 5 seconds). Configure in endpoint settings.
Account limits
Spend limit: Default limit of $80/hour across all resources. Contact support to increase.Billing support
If you believe you’ve been billed incorrectly, contact support, including the following information in your ticket:- Endpoint ID
- Request ID (if applicable)
- Approximate time of the issue