<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Secret123</title>
    <description>The latest articles on DEV Community by Secret123 (@justsecret123).</description>
    <link>https://dev.to/justsecret123</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110061%2Fbe21d514-37c4-43aa-8d14-8e63a7d4f743.png</url>
      <title>DEV Community: Secret123</title>
      <link>https://dev.to/justsecret123</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/justsecret123"/>
    <language>en</language>
    <item>
      <title>Imag-Eval</title>
      <dc:creator>Secret123</dc:creator>
      <pubDate>Fri, 04 Sep 2026 15:46:30 +0000</pubDate>
      <link>https://dev.to/justsecret123/imag-eval-5dn7</link>
      <guid>https://dev.to/justsecret123/imag-eval-5dn7</guid>
      <description>&lt;p&gt;How well do Text-to-Image models actually follow complex instructions? Imag-Eval provides interpretable, skill-based evaluation of T2I instruction following, without error propagation. 📄 Accepted at EMNLP 2026 | ⭐ Check it out &amp;amp; star the repo!&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Justsecret123" rel="noopener noreferrer"&gt;
        Justsecret123
      &lt;/a&gt; / &lt;a href="https://github.com/Justsecret123/Imag-Eval" rel="noopener noreferrer"&gt;
        Imag-Eval
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Imag-Eval [EMNLP 2026] is a skill-based evaluation framework for Text-to-Image (T2I) models that measures instruction-following capabilities across compositional visual reasoning skills, while explicitly controlling for prompt complexity and minimizing error propagation during evaluation. 
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/f83d609947fc061dd3877bbb6e91e6380ea4fbd79a79830714f9c67eccd442d7/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f507974686f6e2d332e31312e342d627269676874677265656e3f7374796c653d666f722d7468652d6261646765266c6f676f3d507974686f6e"&gt;&lt;img src="https://camo.githubusercontent.com/f83d609947fc061dd3877bbb6e91e6380ea4fbd79a79830714f9c67eccd442d7/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f507974686f6e2d332e31312e342d627269676874677265656e3f7374796c653d666f722d7468652d6261646765266c6f676f3d507974686f6e" alt="Static Badge"&gt;&lt;/a&gt; &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/aeb030be2305e22c62bc9ff4996d9f507a0f7cdd6ecfc15d5438ce468e176d6a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f5079546f7263682d322e31302e3025324263753132382d627269676874677265656e3f7374796c653d666f722d7468652d6261646765266c6f676f3d5079746f726368"&gt;&lt;img src="https://camo.githubusercontent.com/aeb030be2305e22c62bc9ff4996d9f507a0f7cdd6ecfc15d5438ce468e176d6a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f5079546f7263682d322e31302e3025324263753132382d627269676874677265656e3f7374796c653d666f722d7468652d6261646765266c6f676f3d5079746f726368" alt="Static Badge"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Imag-Eval [EMNLP 2026]&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Official repository for the paper "&lt;a href="https://arxiv.org/abs/2608.29210" rel="nofollow noopener noreferrer"&gt;Imag-Eval A language-grounded framework for interpretable Text-to-Image instruction following evaluation&lt;/a&gt;". (SEROUIS et al., EMNLP 2026)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Imag-Eval is a skill-based evaluation framework for Text-to-Image (T2I) models that measures instruction-following capabilities across compositional visual reasoning skills, while explicitly controlling for prompt complexity and minimizing error propagation during evaluation; we are rying to shift evaluation paradigms towards more controlled increases in complexity.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2608.29210" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/bde6aaa49e40e0735dccae0b6d51d6b627ca2569a3fc05df0185eb9798fa01a7/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f50617065722d61725869762d6233316231623f7374796c653d666f722d7468652d6261646765266c6f676f3d6172786976266c6f676f436f6c6f723d7768697465" alt="Paper"&gt;&lt;/a&gt; &lt;a href="https://huggingface.co/datasets/JustSecret/Imag-Eval" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/9ef6d1e7ed512af1df5555fd7bdd1dc66162cc675c861515148306482a985a49/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446174617365742d48756767696e67466163652d4646443231453f7374796c653d666f722d7468652d6261646765266c6f676f3d68756767696e6766616365266c6f676f436f6c6f723d626c61636b" alt="Dataset"&gt;&lt;/a&gt; &lt;a href="https://justsecret123.github.io/imag-eval-leaderboard.io" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/e1172b4c110a501fe5c9bdc3b93b46bebb2962d3a8ad61eb90211475259b2de5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6561646572626f6172642d4c6976652d3633363666313f7374796c653d666f722d7468652d6261646765266c6f676f3d6769746875627061676573266c6f676f436f6c6f723d7768697465" alt="Leaderboard"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;To ensure leaderboard integrity and reproducibility, submissions must include the generated images, generation seed, the method used for annotation, and all relevant inference parameters. Reported results will be independently verified, and entries whose reproduced results closely match the submitted scores will be added to the leaderboard.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🚀 Recent News&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;[Sept 2026]&lt;/strong&gt; 🎉 &lt;a href="https://justsecret123.github.io/imag-eval-leaderboard.io" rel="nofollow noopener noreferrer"&gt;The official leaderboard&lt;/a&gt; is now available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;[Aug 2026]&lt;/strong&gt; 🎉 IMAG-EVAL has been accepted to &lt;strong&gt;EMNLP 2026&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;[Aug 2026]&lt;/strong&gt; 📊 Released the first version of the IMAG-EVAL benchmark…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Justsecret123/Imag-Eval" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>deeplearning</category>
      <category>computervision</category>
    </item>
  </channel>
</rss>
