Guides

Evaluating on your data

Evaluate (in the console, with a free account) runs a whole labelled dataset through Wity and compares every answer with the one you expected. This page describes the dataset format.

The flow#

  • Prepare. See the four kinds of question, then download the Excel or JSON template, or paste the prompt on the page into any chat assistant: it asks whether to convert a file you attach or build a dataset with you, and gives you the JSON.
  • Upload. Drop your files, paste JSON, or try the sample. Everything is read in your browser; nothing is uploaded yet.
  • Review. Every problem is listed with how to fix it, and most can be fixed in place.
  • Run. Pick a reasoning mode. Running is paid from your credit at the API's price, only for answered questions. The run continues on our servers if you close the tab.
  • Results. Per type, per file, per question and per row, each metric with its n and next to a baseline. Export as CSV or JSON.

Excel#

One sheet per type: Choice, Score, YesNo and Generate. One row per question; the header row names the columns. The Instructions sheet is not read, and empty rows are skipped.

state
every sheetrequired
The text Wity reads. Up to 32,000 characters. Rows with exactly the same state (in one file) are sent together, up to 32 questions per call.
question
Choice, Score, YesNorequired
The question, in plain words. Generate uses instruction.
options
Choicerequired
2 to 256 options in one cell, one per line (Alt+Enter) or separated by |, each as Label: description.
levels
Scorerequired
2 to 10 levels, lowest first, written like options.
meanings
YesNo
Optional. What yes and no mean, in one cell: Yes: description and No: description, one per line or separated by |. Give both or leave it empty. In JSON this is criteria, with "true" and "false" keys, as the API takes it.
expected_answer
Choice, YesNorequired
The correct option's label. For YesNo: yes or no (y/n, true/false and 1/0 also work).
expected_level
Scorerequired
The correct level's label.
output_format, json_schema, max_tokens
Generate
text or json; for json an optional JSON Schema object; 1 to 512 tokens (default 128).
reference_answer
Generate
Optional. Generate outputs are shown next to it, not graded.
id, notes
every sheet
Optional. Your own id for the row, and notes Wity never sees.

JSON#

The same fields, shaped like the API. A bare list of cases also works. Syntax errors are reported with their line and column.

{
"name": "Support ticket routing",
"cases": [
{
"id": "case-1",
"state": "Customer says they were charged twice this month for the Pro plan and wants the duplicate charge removed.",
"questions": [
{
"id": "case-1-team",
"type": "choice",
"question": "Which team should handle this?",
"options": [
{
"label": "Billing",
"description": "Payments, refunds, invoices"
},
{
"label": "Tech",
"description": "Bugs, outages, errors"
}
],
"expected": "Billing"
},
{
"id": "case-1-urgency",
"type": "score",
"question": "How urgent is this?",
"levels": [
"Low",
"Medium",
"High",
"Critical"
],
"expected": "Medium"
},
{
"id": "case-1-refund",
"type": "noul",
"question": "Is the customer asking for a refund?",
"criteria": {
"true": "Asks for money back, in full or in part",
"false": "Asks only for a fix, a replacement or information"
},
"expected": true
},
{
"id": "case-1-reply",
"type": "generate",
"instruction": "Write a one-line acknowledgement to the customer",
"format": "text",
"reference": "Sorry about the double charge, we're looking into it."
}
]
}
]
}

What is checked#

Errors keep a row from running (you can choose to skip those rows); warnings never do. Among the errors: missing required fields, fewer than 2 options, duplicate labels, an expected answer that is not one of the options (with a one-click “Did you mean”), Yes / No meanings given for only one answer, an invalid JSON schema, and a state over the API's length limit. Warnings include options without descriptions, letter case fixed in an expected answer, duplicate questions, and questions where one expected answer is 80% or more of 20+ rows.

How results are scored#

  • The answer is the most likely option of Wity's distribution. For Yes / No, yes when P(yes) is at least 0.5.
  • Choice: accuracy, balanced accuracy and macro F1, next to the accuracy of always guessing each question's most common expected label.
  • Score: exact-level accuracy, within one level, and mean absolute error in levels, with the same kind of baseline.
  • Generate is not graded. You can mark outputs as good or not good.
  • Questions that failed are left out of the metrics, not counted as wrong.

Confidence is not correctness

Confidence measures how peaked Wity's distribution is. It is uncalibrated and is not a probability of being correct. Results show how accuracy varied with confidence on your data, not a calibration.

Limits#

  • Up to 10 files at once, 5 MB each, and 1,000 questions per evaluation. Text only for now.
  • Up to 2 evaluations running at once per account.
  • New accounts get a free trial: up to 200 questions in all, across your evaluations, in the first 48 hours, paid from the $5 sign-up credit (Compare asks each twice; failed questions don't count). After 48 hours, or once they are used, add credit to keep evaluating. After one top-up there is no question limit.
  • The original files are never stored; the questions, answers and metrics are, until you delete the evaluation.