# Deploy an HTTP container

Source: https://docs.oiy.ai/docs/http-quickstart

Create a service with curl, access its endpoint, and pause and resume compute.



Run a small Python HTTP server using Oiy's catalog runtime image. This walkthrough uses CPU-only resources so you can learn the container lifecycle before choosing a GPU. It still allocates billable compute and storage when accepted. It is a learning server, not a production inference server.

## 1. Prepare your account [#1-prepare-your-account]

You need a verified account, sufficient balance, an enabled placement with compatible capacity, and a personal [API key](/docs/api/authentication). Use a local terminal with `curl`, `jq`, and `uuidgen`. Supply `OIY_API_KEY` through your environment or secret store, then set the API origin:

```sh
export OIY_API_URL="https://oiy-ai-server.xsun.workers.dev"
curl --fail-with-body --silent --show-error \
  "$OIY_API_URL/api/catalog" > catalog.json
jq '{regions, templates, payment}' catalog.json
```

Review the Standard placement and current prices. A new account starts with zero credit; if top-ups are unavailable, an unfunded account cannot proceed. The catalog is not a capacity reservation. Read [preview availability](/docs/preview) if your deployment differs from this reference.

## 2. Save your container configuration [#2-save-your-container-configuration]

Save the following as `service.json` in a new working directory. Its image is the digest-pinned runtime used by the current built-in templates, including Python and the required runtime tools. You can compare it with `templates` in the catalog above before proceeding.

```json
{
  "name": "hello-container",
  "region": "standard",
  "geography": "auto",
  "image": "oiy-ai-pytorch-images.xsun.workers.dev/oiy-ai-pytorch@sha256:75e3759275db19bdf2cf34b7dd59e2fc8bbf381faf91a194f05224a0f8d62302",
  "command": ["python", "-m", "http.server", "8080", "--bind", "0.0.0.0", "--directory", "/workspace"],
  "resources": { "gpuCount": 0, "cpu": 2, "memoryGb": 4 },
  "httpPort": 8080,
  "volumeGb": 20,
  "volumeMode": "service",
  "sleepAfterMinutes": 15
}
```

The command serves `/workspace` on port `8080`. Files in that mount survive sleep and restart. Here the volume belongs to the service: deleting the service also deletes its workspace. Use [independent storage](/docs/storage) when files must outlive a service.

## 3. Create once, then inspect [#3-create-once-then-inspect]

The next request creates paid resources. Review your configuration and balance first. Generate an operation ID once; keep it unchanged when retrying this exact request.

```sh
CREATE_REQUEST_ID="$(uuidgen)"
curl --fail-with-body --silent --show-error \
  "$OIY_API_URL/api/services" \
  -H "Authorization: Bearer $OIY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $CREATE_REQUEST_ID" \
  --data-binary @service.json > created.json
```

Continue only after curl succeeds. Creation returns a bare service object; save its ID:

```sh
SERVICE_ID="$(jq -er '.id' created.json)"
curl --fail-with-body --silent --show-error \
  "$OIY_API_URL/api/services/$SERVICE_ID" \
  -H "Authorization: Bearer $OIY_API_KEY" > service-state.json
jq '.service | {id, status, allocated, error, endpoint}' service-state.json
```

Repeat the **GET** request until the service is `running`. `queued` and `starting` are intermediate states. If it enters `error`, inspect [events and logs](/docs/services/logs). A failed or interrupted create request can have an uncertain outcome: inspect your account before retrying with the same operation ID. Do not generate another ID just because the first response was lost.

## 4. Open the authenticated endpoint [#4-open-the-authenticated-endpoint]

Use the endpoint returned by the service:

```sh
OIY_SERVICE_ENDPOINT="$(jq -er '.service.endpoint' service-state.json)"
curl --fail-with-body --silent --show-error \
  "$OIY_SERVICE_ENDPOINT" \
  -H "Authorization: Bearer $OIY_API_KEY"
```

You should receive an HTML directory listing of the workspace, or its `index.html` if one exists. A `running` service can still be initializing its application. For `503`, follow `Retry-After` and check the logs. Do not place your API key in a URL. [Browser launch](/docs/services/endpoints) is a separate console flow.

## 5. Pause and confirm release [#5-pause-and-confirm-release]

```sh
STOP_REQUEST_ID="$(uuidgen)"
curl --fail-with-body --silent --show-error \
  "$OIY_API_URL/api/services/$SERVICE_ID/actions" \
  -H "Authorization: Bearer $OIY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $STOP_REQUEST_ID" \
  --data '{"action":"stop"}'
```

Repeat the service GET from step 3. Wait for `status: "sleeping"` and `allocated: false`; a stop acknowledgement alone does not prove compute has been released. Compute charges end after confirmed release. Retained storage remains billable. Avoid requesting the HTTP endpoint while checking sleep: that request can wake the service.

## 6. Resume when needed [#6-resume-when-needed]

```sh
START_REQUEST_ID="$(uuidgen)"
curl --fail-with-body --silent --show-error \
  "$OIY_API_URL/api/services/$SERVICE_ID/actions" \
  -H "Authorization: Bearer $OIY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $START_REQUEST_ID" \
  --data '{"action":"start"}'
```

Resuming allocates compute again, subject to balance and capacity. Wait for `running`, then access the endpoint. Files in `/workspace` remain; the Python process starts again. Pause again when finished. To stop storage charges too, back up any files and delete this service through the console; its service-owned workspace will be removed.

## Move to a GPU workload [#move-to-a-gpu-workload]

For a guided GPU calculation with the built-in runtime, continue with [Run a GPU notebook](/docs/gpu-notebook).

Choose an available [GPU profile](/docs/compute), replace the learning server with your application command, and use the matching image dependencies. A GPU resource configuration alone does not turn this file server into an AI application. Long jobs should use [activity protection](/docs/services/sleep).

`sleepAfterMinutes: 15` configures automatic idle sleep. It does not mean an active connection or protected task is killed at fifteen minutes. `gpuCount: 0` means CPU-only compute; **sleeping with no allocation** is scale-to-zero.
