Guides
Evaluating on your data
Evaluate (in the console, with a free account) runs a whole labelled dataset through Wity and compares every answer with the one you expected. This page describes the dataset format.
The flow#
- Prepare. See the four kinds of question, then download the Excel or JSON template, or paste the prompt on the page into any chat assistant: it asks whether to convert a file you attach or build a dataset with you, and gives you the JSON.
- Upload. Drop your files, paste JSON, or try the sample. Everything is read in your browser; nothing is uploaded yet.
- Review. Every problem is listed with how to fix it, and most can be fixed in place.
- Run. Pick a reasoning mode. Running is paid from your credit at the API's price, only for answered questions. The run continues on our servers if you close the tab.
- Results. Per type, per file, per question and per row, each metric with its n and next to a baseline. Export as CSV or JSON.
Excel#
One sheet per type: Choice, Score, YesNo and Generate. One row per question; the header row names the columns. The Instructions sheet is not read, and empty rows are skipped.
instruction.|, each as Label: description.Yes: description and No: description, one per line or separated by |. Give both or leave it empty. In JSON this is criteria, with "true" and "false" keys, as the API takes it.text or json; for json an optional JSON Schema object; 1 to 512 tokens (default 128).JSON#
The same fields, shaped like the API. A bare list of cases also works. Syntax errors are reported with their line and column.
{"name": "Support ticket routing","cases": [{"id": "case-1","state": "Customer says they were charged twice this month for the Pro plan and wants the duplicate charge removed.","questions": [{"id": "case-1-team","type": "choice","question": "Which team should handle this?","options": [{"label": "Billing","description": "Payments, refunds, invoices"},{"label": "Tech","description": "Bugs, outages, errors"}],"expected": "Billing"},{"id": "case-1-urgency","type": "score","question": "How urgent is this?","levels": ["Low","Medium","High","Critical"],"expected": "Medium"},{"id": "case-1-refund","type": "noul","question": "Is the customer asking for a refund?","criteria": {"true": "Asks for money back, in full or in part","false": "Asks only for a fix, a replacement or information"},"expected": true},{"id": "case-1-reply","type": "generate","instruction": "Write a one-line acknowledgement to the customer","format": "text","reference": "Sorry about the double charge, we're looking into it."}]}]}
What is checked#
Errors keep a row from running (you can choose to skip those rows); warnings never do. Among the errors: missing required fields, fewer than 2 options, duplicate labels, an expected answer that is not one of the options (with a one-click “Did you mean”), Yes / No meanings given for only one answer, an invalid JSON schema, and a state over the API's length limit. Warnings include options without descriptions, letter case fixed in an expected answer, duplicate questions, and questions where one expected answer is 80% or more of 20+ rows.
How results are scored#
- The answer is the most likely option of Wity's distribution. For Yes / No, yes when P(yes) is at least 0.5.
- Choice: accuracy, balanced accuracy and macro F1, next to the accuracy of always guessing each question's most common expected label.
- Score: exact-level accuracy, within one level, and mean absolute error in levels, with the same kind of baseline.
- Generate is not graded. You can mark outputs as good or not good.
- Questions that failed are left out of the metrics, not counted as wrong.
Confidence is not correctness
Limits#
- Up to 10 files at once, 5 MB each, and 1,000 questions per evaluation. Text only for now.
- Up to 2 evaluations running at once per account.
- New accounts get a free trial: up to 200 questions in all, across your evaluations, in the first 48 hours, paid from the $5 sign-up credit (Compare asks each twice; failed questions don't count). After 48 hours, or once they are used, add credit to keep evaluating. After one top-up there is no question limit.
- The original files are never stored; the questions, answers and metrics are, until you delete the evaluation.
