Skip to content
For the complete documentation index, see llms.txt. Markdown versions of documentation pages are available by appending .md to the page URL.
Primary navigation

Get an eval

client.evals.retrieve(stringevalID, RequestOptionsoptions?): EvalRetrieveResponse { id, created_at, data_source_config, 4 more }
GET/evals/{eval_id}

Get an evaluation by ID.

ParametersExpand Collapse
evalID: string
ReturnsExpand Collapse
EvalRetrieveResponse { id, created_at, data_source_config, 4 more }

An Eval object with a data source config and testing criteria. An Eval represents a task to be done for your LLM integration. Like:

  • Improve the quality of my chatbot
  • See how well my chatbot handles customer support
  • Check if o4-mini is better at my usecase than gpt-4o
id: string

Unique identifier for the evaluation.

created_at: number

The Unix timestamp (in seconds) for when the eval was created.

formatunixtime
data_source_config: EvalCustomDataSourceConfig { schema, type } | Logs { schema, type, metadata } | EvalStoredCompletionsDataSourceConfig { schema, type, metadata }

Configuration of data sources used in runs of the evaluation.

metadata: Metadata | null

Set of 16 key-value pairs that can be attached to an object. This can be useful for storing additional information about the object in a structured format, and querying for objects via API or the dashboard.

Keys are strings with a maximum length of 64 characters. Values are strings with a maximum length of 512 characters.

name: string

The name of the evaluation.

object: "eval"

The object type.

testing_criteria: Array<LabelModelGrader { input, labels, model, 3 more } | StringCheckGrader { input, name, operation, 2 more } | EvalGraderTextSimilarity { pass_threshold } | 2 more>

A list of testing criteria.

Get an eval

import OpenAI from "openai";

const openai = new OpenAI();

const evalObj = await openai.evals.retrieve("eval_67abd54d9b0081909a86353f6fb9317a");
console.log(evalObj);
{
  "object": "eval",
  "id": "eval_67abd54d9b0081909a86353f6fb9317a",
  "data_source_config": {
    "type": "custom",
    "schema": {
      "type": "object",
      "properties": {
        "item": {
          "type": "object",
          "properties": {
            "input": {
              "type": "string"
            },
            "ground_truth": {
              "type": "string"
            }
          },
          "required": [
            "input",
            "ground_truth"
          ]
        }
      },
      "required": [
        "item"
      ]
    }
  },
  "testing_criteria": [
    {
      "name": "String check",
      "id": "String check-2eaf2d8d-d649-4335-8148-9535a7ca73c2",
      "type": "string_check",
      "input": "{{item.input}}",
      "reference": "{{item.ground_truth}}",
      "operation": "eq"
    }
  ],
  "name": "External Data Eval",
  "created_at": 1739314509,
  "metadata": {},
}
Returns Examples
{
  "object": "eval",
  "id": "eval_67abd54d9b0081909a86353f6fb9317a",
  "data_source_config": {
    "type": "custom",
    "schema": {
      "type": "object",
      "properties": {
        "item": {
          "type": "object",
          "properties": {
            "input": {
              "type": "string"
            },
            "ground_truth": {
              "type": "string"
            }
          },
          "required": [
            "input",
            "ground_truth"
          ]
        }
      },
      "required": [
        "item"
      ]
    }
  },
  "testing_criteria": [
    {
      "name": "String check",
      "id": "String check-2eaf2d8d-d649-4335-8148-9535a7ca73c2",
      "type": "string_check",
      "input": "{{item.input}}",
      "reference": "{{item.ground_truth}}",
      "operation": "eq"
    }
  ],
  "name": "External Data Eval",
  "created_at": 1739314509,
  "metadata": {},
}