<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://chrishounwanou.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://chrishounwanou.github.io/" rel="alternate" type="text/html" /><updated>2026-10-02T11:14:27-07:00</updated><id>https://chrishounwanou.github.io/feed.xml</id><title type="html">Christophe D. Hounwanou</title><subtitle>Ph.D Researcher in Computer Science @ Mila/IID/ULaval</subtitle><author><name>Christophe D. Hounwanou</name><email>chrishounwanou@gmail.com</email></author><entry><title type="html">Tiny-Aya-Water Blind Spot Evaluation Dataset</title><link href="https://chrishounwanou.github.io/others/tiny-aya-water-blind-spots/" rel="alternate" type="text/html" title="Tiny-Aya-Water Blind Spot Evaluation Dataset" /><published>2026-02-20T00:00:00-08:00</published><updated>2026-02-20T00:00:00-08:00</updated><id>https://chrishounwanou.github.io/others/other1</id><content type="html" xml:base="https://chrishounwanou.github.io/others/tiny-aya-water-blind-spots/"><![CDATA[<p>This dataset contains <strong>10 manually constructed blind spot cases</strong> designed to evaluate reasoning, instruction following, and factual robustness of the model:</p>

<p><strong>CohereLabs/tiny-aya-water</strong></p>

<p>The goal of this dataset is to identify <strong>systematic weaknesses and failure modes</strong> of the model under controlled experimental conditions.</p>

<h1 id="dataset-link">Dataset link</h1>
<p><a href="https://huggingface.co/datasets/chrishounwanou/tiny-aya-water-blind-spots">Hugging Face Dataset</a></p>

<h1 id="dataset-characteristics">Dataset characteristics</h1>
<ul>
  <li>Small evaluation dataset (10 cases)</li>
  <li>Designed for <strong>LLM robustness testing</strong></li>
  <li>Focus on <strong>reasoning failures and blind spots</strong></li>
  <li>Useful for controlled experimental analysis</li>
</ul>

<h1 id="potential-use-cases">Potential use cases</h1>
<ul>
  <li>Evaluation of <strong>large language models</strong></li>
  <li>Testing <strong>instruction-following robustness</strong></li>
  <li>Studying <strong>failure modes in generative AI systems</strong></li>
</ul>]]></content><author><name>Christophe D. Hounwanou</name><email>chrishounwanou@gmail.com</email></author><category term="dataset" /><category term="llm-evaluation" /><category term="robustness" /><category term="ai-safety" /><summary type="html"><![CDATA[This dataset contains 10 manually constructed blind spot cases designed to evaluate reasoning, instruction following, and factual robustness of the model:]]></summary></entry></feed>