Tiny-Aya-Water Blind Spot Evaluation Dataset

less than 1 minute read

Published:

This dataset contains 10 manually constructed blind spot cases designed to evaluate reasoning, instruction following, and factual robustness of the model:

CohereLabs/tiny-aya-water

The goal of this dataset is to identify systematic weaknesses and failure modes of the model under controlled experimental conditions.

Dataset link

Hugging Face Dataset

Dataset characteristics

  • Small evaluation dataset (10 cases)
  • Designed for LLM robustness testing
  • Focus on reasoning failures and blind spots
  • Useful for controlled experimental analysis

Potential use cases

  • Evaluation of large language models
  • Testing instruction-following robustness
  • Studying failure modes in generative AI systems