Skip to main content
Everything the console’s Training page does is available from code, with the same API key you use for inference. A job started here shows up there, and the other way round. The examples use my-account and your-base-model where your own values go: your account id is the one in the console’s address bar, and the bases you can train are listed on the Training page.

Before you start

  • A key that can train. Any key of a member with the Developer role or above works. A key created with operation scopes needs training (for jobs) and datasets (for uploads) among them. A key only ever reaches the account it was created in.
  • The data processing addendum, if your job runs outside the platform. Some bases are trained by an outside service, which means handing it your dataset. The estimate’s egress.execution.thirdParty tells you which case a job is. For those, your organization accepts the addendum once, in the console; until it does, creating the job returns AGREEMENT_REQUIRED — and none of your data leaves.

Dataset format

A dataset is one JSON Lines file: UTF-8, one JSON object per line, no line over 8 MiB, at most 1 GiB in total. Supervised fine-tuning learns from conversations. The last message must be the assistant’s — that turn is what the model learns to produce; everything before it is context.
Preference training (DPO) learns from a prompt and two answers. The conversation stops at the user’s turn; the preferred and rejected answers sit beside it.
Validation checks that the file is JSON Lines, not that its rows have the shape a kind of training needs — the same file can be used for different jobs. Rows a job cannot read are skipped and counted in the job’s progress.skippedExamples; if no row can be read, the job fails with TRAINING_TOKENIZE_FAILED before anything is billed.

1. Upload the dataset

Three calls: create the dataset, PUT the file to a signed URL, then ask for it to be validated. Your file goes straight to storage; it never passes through the API.
The same steps with cURL — the upload URL, then the PUT:
Then validation, which runs in the background:
Poll the dataset until its state is DATASET_STATE_READY. If it ends in DATASET_STATE_FAILED, the version’s validation says which line was the first bad one and why — never what was on it. To correct a dataset, upload again with the same three calls: the corrected file becomes a new version. Jobs are pinned to the version they started with, so a correction never changes what a finished job trained on.

2. Check the cost

The estimate takes the same body as the job, creates nothing, and shows its working — examples, tokens per example, epochs and rate — along with whether the job would be allowed to run.
You are billed for the tokens actually trained on, which the estimate approximates.

3. Start the job

The 201 means the job has been accepted (JOB_STATE_QUEUED), not that it has started. Send an Idempotency-Key and a retried request returns the same job instead of starting a second one. For preference training, the body is the same and the collection is dpoJobs; config.dpoBeta sets how far the model may move from the base (default 0.1).

4. Follow it

progress carries the epoch, the tokens processed so far and the last points of the training curve; the metrics endpoint returns the whole curve. A job that is queued or running can be cancelled. When the job succeeds, outputModel names the new model in your account. It is listed under Custom models in the console alongside your other models.

5. Serve it

A fine-tuned model is served on a dedicated deployment: open it under Custom models in the console and deploy it. Once the deployment is ready, call it through the inference API with the deployment’s resource name as model — not the model’s:
The model’s own name (accounts/my-account/models/support-lora-v1) identifies what was trained; it is not something you can send requests to, and doing so returns MODEL_NOT_FOUND. A dedicated deployment is billed for as long as it runs, whether or not it is serving requests — delete it when you are done.