Tiny-Aya-Water Blind Spot Evaluation Dataset
Published:
This dataset contains 10 manually constructed blind spot cases designed to evaluate reasoning, instruction following, and factual robustness of the model:
CohereLabs/tiny-aya-water
The goal of this dataset is to identify systematic weaknesses and failure modes of the model under controlled experimental conditions.
Dataset link
Dataset characteristics
- Small evaluation dataset (10 cases)
- Designed for LLM robustness testing
- Focus on reasoning failures and blind spots
- Useful for controlled experimental analysis
Potential use cases
- Evaluation of large language models
- Testing instruction-following robustness
- Studying failure modes in generative AI systems
