API reference
Everything you need to call the inference API: authentication, endpoints, payloads, and errors.
Authentication
Send your API key in the Authorization header as a Bearer token. Keys are environment-scoped: vs_live_… keys hit production, vs_test_… keys hit development and staging. Create and rotate keys on the API keys page.
Quickstart
Run detection on an image in one request
bash
curl -X POST https://api.visionserve.dev/v1/detect \
-H "Authorization: Bearer vs_live_…" \
-H "Content-Type: multipart/form-data" \
-F "image=@street.jpg" \
-F "confidence_threshold=0.25"Endpoints
Base URL: https://api.visionserve.dev
| Method | Path | Description | Required scope |
|---|---|---|---|
| POST | /v1/detect | Run object detection on an image | inference:run |
| POST | /v1/ocr | Extract text from an image | inference:run |
| POST | /v1/classify | Classify an image into labels | inference:run |
| POST | /v1/jobs | Submit an asynchronous inference job | jobs:write |
| GET | /v1/jobs/{id} | Poll job status and progress | jobs:read |
| POST | /v1/batches | Submit a batch of files for processing | batches:write |
| GET | /v1/batches/{id} | Poll batch progress | jobs:read |
| GET | /v1/results/{id} | Fetch a stored inference result | results:read |
| GET | /v1/models | List registered models | models:read |
| GET | /v1/health | Liveness and component status | — |
Error codes
All errors return a JSON body with code, message, and request_id
| Code | HTTP | Description |
|---|---|---|
| INVALID_IMAGE | 400 | The request body could not be decoded as a supported image. |
| UNSUPPORTED_IMAGE_FORMAT | 415 | Format not in the workspace allowlist (JPEG, PNG, WebP, TIFF, HEIC). |
| IMAGE_TOO_LARGE | 413 | Upload exceeds the configured maximum file size. |
| PIXEL_LIMIT_EXCEEDED | 413 | Decoded pixel count exceeds the decompression guard. |
| MODEL_NOT_FOUND | 404 | No model with this ID exists in the registry. |
| MODEL_UNAVAILABLE | 503 | Model failed health checks or is not loaded in this environment. |
| MODEL_LOAD_FAILED | 500 | The runtime session could not be created from the artifact. |
| INFERENCE_TIMEOUT | 504 | Inference exceeded the configured job timeout. |
| QUEUE_FULL | 429 | The job queue is at capacity; retry with backoff. |
| RATE_LIMITED | 429 | The API key exceeded its per-minute rate limit. |
| AUTHENTICATION_REQUIRED | 401 | Missing or invalid API key. |
| PERMISSION_DENIED | 403 | The key lacks the required scope for this operation. |
| STORAGE_ERROR | 500 | Object storage read or write failed. |
| INTERNAL_ERROR | 500 | Unexpected server error; contact support with the request ID. |