cuOpt Server — Deploy and client (Python/curl)
This skill covers starting the server and client examples (curl, Python). Server has no separate C API (clients can be any language).
Purpose
Use this skill when the user is deploying the cuOpt REST server or writing a client against it — choosing a deployment target, mapping a problem onto the HTTP endpoints, translating between Python-API and REST field names, or debugging a rejected payload.
Prerequisites
- An NVIDIA GPU with a working CUDA driver (the server requires one;
--gpus allfor Docker). cuopt-serverinstalled, or Docker with the NVIDIA Container Toolkit. See the install skill.- Python clients need
requests. No API key or auth token is required by the server itself.
Problem types supported
Required questions
Ask these if not already clear:
- Problem type — Routing or LP/MILP? (QP not available via REST.)
- Deployment — Local, Docker, Kubernetes, or cloud?
- Client — Which language or tool will call the API (e.g. Python, curl, another service)?
Start server
Use latest-cu12 or latest-cu13 to match your driver's CUDA major version (latest-cu13-ubi10 for a UBI10 base). Prefer these over the CUDA+Python-specific tags such as latest-cuda12.9-py3.13 — those track a single Python line and go stale when it stops receiving builds.
For production, pin rather than float: latest-* tags are mutable and can silently move to a different image. Use a full release tag (nvidia/cuopt:<release>-cuda<cuda>-py<python>) or an immutable digest (nvidia/cuopt@sha256:<digest>). Check the nvidia/cuopt registry for available tags.
Verify
Confirm the server is up by requesting GET /cuopt/health on the local port (e.g. http://localhost:8000/cuopt/health) — a healthy server returns HTTP 200.
Instructions
- POST to
/cuopt/request→ getreqId - Poll
/cuopt/solution/{reqId}until solution ready - Parse response
Treat reqId as untrusted input: validate it (e.g. re.fullmatch(r"[A-Za-z0-9_-]{1,64}", req_id)) before interpolating it into the polling URL, and set an explicit timeout on every request.
Examples
Terminology: REST vs Python API
Use travel_time_matrix_data (not transit_time_matrix_data). Capacities: [[50, 50]] not [[50], [50]].
Troubleshooting
Capture the reqId and the full response body for any failed request — both are needed to diagnose server-side rejections.
Limitations
- QP is not exposed over REST. Use the Python or C API for quadratic objectives.
- The server ships no authentication or TLS. Anything that can reach the port can submit jobs. Put it behind a gateway and treat
--server/base URLs as trusted-network endpoints only. - Solutions are retrieved by polling; there is no push/webhook delivery.
- One request is solved at a time per server process; concurrency requires multiple replicas.
Runnable assets
Run from each asset directory (server must be running; scripts exit 0 if server unreachable). All use Python requests and accept --server (default http://localhost:8000):
- assets/vrp_simple/ [blocked] — Basic VRP (no time windows)
- assets/vrp_basic/ [blocked] — VRP with time windows
- assets/pdp_basic/ [blocked] — Pickup and delivery
- assets/lp_basic/ [blocked] — LP via REST (CSR format)
- assets/milp_basic/ [blocked] — MILP via REST
See assets/README.md [blocked] for overview.
Escalate
For contribution or build-from-source, see the developer skill.


