Search pages and blog articles, or ask Pilla.
Pulling account data to ground an answer or a decision.
Graded by a fixed judge scoring every case 0, 0.5 or 1 against a written rubric. The score is the mean of those grades, so a case can land half-right — it is an average grade, not a pass rate.
The cases behind that number, grouped by what they are trying to catch. Open one to read the actual prompt Pilla was given and the answer she gave back.
Prompting an action
12 test cases
Audience and voice
12 test cases
Basics
3 test cases
Data accuracy
15 test cases
Edge cases and safety
12 test cases
Format instructions
14 test cases
Grounding
3 test cases
Deep grounding
14 test cases
Grounding traps
10 test cases
Harmful requests
12 test cases
Prompt injection
14 test cases
Malformed input
16 test cases
Contradictory instructions
14 test cases
One-way instructions
3 test cases
One-way instructions, extended
8 test cases
Scope and time
14 test cases
Sensitive personal data
10 test cases
Slack context
2 test cases
Tone
8 test cases
Optional cookies
We use optional analytics cookies to understand how Pilla is used and LinkedIn advertising cookies to measure campaigns and show relevant ads. LinkedIn does not load unless you agree. See our Privacy Policy.