MoltinMoltin Docs

Datasets

Datasets are collections of evaluation data used to test agents. Manage them from Build → Datasets (/agents/datasets); they exist at the workspace level only (there is no per-agent datasets page).

🎬Opening Datasets from the Build sidebar.Watch
Opening Datasets from the Build sidebar.

The page lays out a title, a search box, and a Create Dataset button across the header.

🎬The Datasets page structure: title, search, and Create Dataset button.Watch
The Datasets page structure: title, search, and Create Dataset button.

Controls

The Datasets page header has a Create Dataset button to start a new dataset.

🎬Use Create Dataset to start a new dataset.Watch
Use Create Dataset to start a new dataset.

It also offers Filter By and Sort By controls alongside a search box (Search datasets…) to narrow the list.

🎬Filter By and Sort By controls for narrowing the datasets list.Watch
Filter By and Sort By controls for narrowing the datasets list.
🎬The Search datasets… box filters the list by keyword.Watch
The Search datasets… box filters the list by keyword.

A dataset count label is shown in the form "X of Y datasets".

🎬The dataset count label reads 'X of Y datasets'.Watch
The dataset count label reads 'X of Y datasets'.

The datasets table

The table has Name, Agent, Last Evaluation, and Updated On columns. Each row shows the dataset name, the associated agent (with its avatar), and the last-evaluation / updated timestamps.

🎬The Datasets table: Name, Agent, Last Evaluation, and Updated On columns with the associated agent per row.Watch
The Datasets table: Name, Agent, Last Evaluation, and Updated On columns with the associated agent per row.
🎬Each dataset row shows the associated agent's name and avatar.Watch
Each dataset row shows the associated agent's name and avatar.

The table shows at least one dataset row whenever the workspace has datasets.

🎬The table lists each dataset as a row.Watch
The table lists each dataset as a row.

Datasets power evaluations

Pair a dataset with an agent to run evaluations and track results over time — the Last Evaluation column reflects the most recent run.