See
Reasoning tokens and max_tokens
.
Warm-up times usually range from a few seconds to a few minutes, depending on the model and provider.
If you encounter persistent no-content issues, consider implementing a simple retry mechanism or trying again with a different provider or model that has more recent activity.
In some cases, you may still be charged for the prompt processing cost by the upstream provider, even if no content is generated.
​
Streaming error formats
When using streaming mode (
stream: true
), errors are handled differently depending on when they occur:
​
Pre-stream errors
Errors that occur before any tokens are sent follow the standard error format above, with appropriate HTTP status codes.
