> ## Documentation Index
> Fetch the complete documentation index at: https://docs.compute.prentis.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Fine-tuning

> Train a model on your own data from code: upload a dataset, start a job, follow it to the end.

Everything the console's **Training** page does is available from code, with the same API key
you use for inference. A job started here shows up there, and the other way round.

The examples use `my-account` and `your-base-model` where your own values go: your account id
is the one in the console's address bar, and the bases you can train are listed on the
**Training** page.

## Before you start

* **A key that can train.** Any key of a member with the Developer role or above works. A key
  created with operation scopes needs `training` (for jobs) and `datasets` (for uploads) among
  them. A key only ever reaches the account it was created in.
* **The data processing addendum, if your job runs outside the platform.** Some bases are
  trained by an outside service, which means handing it your dataset. The estimate's
  `egress.execution.thirdParty` tells you which case a job is. For those, your organization
  accepts the addendum once, in the console; until it does, creating the job returns
  `AGREEMENT_REQUIRED` -- and none of your data leaves.

## Dataset format

A dataset is one [JSON Lines](https://jsonlines.org) file: UTF-8, one JSON object per line,
no line over 8 MiB, at most 1 GiB in total.

**Supervised fine-tuning** learns from conversations. The last message must be the assistant's
\-- that turn is what the model learns to produce; everything before it is context.

```json theme={null}
{"messages": [{"role": "system", "content": "You are a careful assistant."}, {"role": "user", "content": "What is the capital of Australia?"}, {"role": "assistant", "content": "Canberra."}]}
```

**Preference training (DPO)** learns from a prompt and two answers. The conversation stops at
the user's turn; the preferred and rejected answers sit beside it.

```json theme={null}
{"messages": [{"role": "user", "content": "What is the capital of Australia?"}], "chosen": {"role": "assistant", "content": "Canberra."}, "rejected": {"role": "assistant", "content": "Sydney, which is also its largest city."}}
```

<Note>
  Validation checks that the file is JSON Lines, not that its rows have the shape a kind of
  training needs -- the same file can be used for different jobs. Rows a job cannot read are
  skipped and counted in the job's `progress.skippedExamples`; if no row can be read, the job
  fails with `TRAINING_TOKENIZE_FAILED` before anything is billed.
</Note>

## 1. Upload the dataset

Three calls: create the dataset, PUT the file to a signed URL, then ask for it to be
validated. Your file goes straight to storage; it never passes through the API.

```python theme={null}
import os
import time
import requests

BASE = "https://compute.prentis.ai/v1/accounts/my-account"
HEADERS = {"Authorization": f"Bearer {os.environ['PRENTIS_API_KEY']}"}
path = "train.jsonl"

# 1. Create the dataset.
requests.post(f"{BASE}/datasets", params={"datasetId": "support-chats"},
              headers=HEADERS, json={"displayName": "Support chats"}).raise_for_status()

# 2. Ask for an upload URL, then PUT the file straight to it.
up = requests.post(f"{BASE}/datasets/support-chats:getUploadEndpoint", headers=HEADERS,
                   json={"filenameToSize": {path: str(os.path.getsize(path))}})
up.raise_for_status()
with open(path, "rb") as f:
    requests.put(up.json()["filenameToSignedUrl"][path], data=f).raise_for_status()

# 3. Validate, and wait for the verdict.
requests.post(f"{BASE}/datasets/support-chats:validateUpload", headers=HEADERS,
              json={"uploadSession": up.json()["uploadSession"]["name"]}).raise_for_status()
while (ds := requests.get(f"{BASE}/datasets/support-chats", headers=HEADERS).json())["state"] \
        == "DATASET_STATE_VALIDATING":
    time.sleep(5)
print(ds["state"])
```

The same steps with cURL -- the upload URL, then the PUT:

```bash theme={null}
curl https://compute.prentis.ai/v1/accounts/my-account/datasets/support-chats:getUploadEndpoint \
  -H "Authorization: Bearer $PRENTIS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"filenameToSize": {"train.jsonl": "2142208"}}'

# Then upload the file to the URL in the response -- no API key on this request:
curl -X PUT --upload-file train.jsonl "<filenameToSignedUrl[train.jsonl]>"
```

Then validation, which runs in the background:

```bash theme={null}
curl https://compute.prentis.ai/v1/accounts/my-account/datasets/support-chats:validateUpload \
  -H "Authorization: Bearer $PRENTIS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"uploadSession": "accounts/my-account/uploadSessions/01K7AWZ8C4E6G8J0L2N4Q6S8U0"}'
```

Poll the dataset until its `state` is `DATASET_STATE_READY`. If it ends in
`DATASET_STATE_FAILED`, the version's `validation` says which line was the first bad one and
why -- never what was on it.

To correct a dataset, upload again with the same three calls: the corrected file becomes a
new version. Jobs are pinned to the version they started with, so a correction never changes
what a finished job trained on.

## 2. Check the cost

The estimate takes the same body as the job, creates nothing, and shows its working --
examples, tokens per example, epochs and rate -- along with whether the job would be allowed
to run.

```bash theme={null}
curl https://compute.prentis.ai/v1/accounts/my-account/supervisedFineTuningJobs:estimateCost \
  -H "Authorization: Bearer $PRENTIS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "baseModel": "accounts/maas/models/your-base-model",
    "inputDatasetVersion": "accounts/my-account/datasets/support-chats",
    "config": { "outputModelId": "support-lora-v1", "epochs": 2 }
  }'
```

You are billed for the tokens actually trained on, which the estimate approximates.

## 3. Start the job

<CodeGroup>
  ```bash theme={null}
  curl https://compute.prentis.ai/v1/accounts/my-account/supervisedFineTuningJobs \
    -H "Authorization: Bearer $PRENTIS_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: support-lora-v1" \
    -d '{
      "baseModel": "accounts/maas/models/your-base-model",
      "inputDatasetVersion": "accounts/my-account/datasets/support-chats",
      "config": {
        "outputModelId": "support-lora-v1",
        "epochs": 2,
        "learningRate": "1e-4",
        "loraRank": 16
      }
    }'
  ```

  ```python theme={null}
  import os
  import requests

  BASE = "https://compute.prentis.ai/v1/accounts/my-account"
  HEADERS = {"Authorization": f"Bearer {os.environ['PRENTIS_API_KEY']}"}

  job = requests.post(
      f"{BASE}/supervisedFineTuningJobs",
      headers={**HEADERS, "Idempotency-Key": "support-lora-v1"},
      json={
          "baseModel": "accounts/maas/models/your-base-model",
          "inputDatasetVersion": "accounts/my-account/datasets/support-chats",
          "config": {"outputModelId": "support-lora-v1", "epochs": 2, "learningRate": "1e-4"},
      },
  )
  job.raise_for_status()
  print(job.json()["name"], job.json()["state"])
  ```
</CodeGroup>

The `201` means the job has been accepted (`JOB_STATE_QUEUED`), not that it has started.
Send an `Idempotency-Key` and a retried request returns the same job instead of starting a
second one.

For preference training, the body is the same and the collection is
[`dpoJobs`](/api-reference/create-dpo-job); `config.dpoBeta` sets how far the model may move
from the base (default `0.1`).

## 4. Follow it

```python theme={null}
import os
import time
import requests

url = "https://compute.prentis.ai/v1/accounts/my-account/supervisedFineTuningJobs/01K7AX3V9Q2M8T4R6Y0B1C5D7E"
headers = {"Authorization": f"Bearer {os.environ['PRENTIS_API_KEY']}"}

done = {"JOB_STATE_SUCCEEDED", "JOB_STATE_FAILED", "JOB_STATE_CANCELLED", "JOB_STATE_EXPIRED"}
while True:
    job = requests.get(url, headers=headers).json()
    print(job["state"], job.get("progress", {}).get("percent", 0), "%")
    if job["state"] in done:
        break
    time.sleep(30)
print("model:", job.get("outputModel"))
```

`progress` carries the epoch, the tokens processed so far and the last points of the training
curve; [the metrics endpoint](/api-reference/get-supervised-fine-tuning-job-metrics) returns
the whole curve. A job that is queued or running can be
[cancelled](/api-reference/cancel-supervised-fine-tuning-job).

When the job succeeds, `outputModel` names the new model in your account. It is listed under
**Custom models** in the console alongside your other models.

## 5. Serve it

A fine-tuned model is served on a **dedicated deployment**: open it under **Custom models** in
the console and deploy it. Once the deployment is ready, call it through the inference API
with the deployment's resource name as `model` -- not the model's:

```json theme={null}
{"model": "accounts/my-account/deployments/support-lora-v1", "messages": [{"role": "user", "content": "Hello"}]}
```

The model's own name (`accounts/my-account/models/support-lora-v1`) identifies what was
trained; it is not something you can send requests to, and doing so returns
`MODEL_NOT_FOUND`. A dedicated deployment is billed for as long as it runs, whether or not it
is serving requests -- delete it when you are done.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.