<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: tutorial</title>
    <description>The latest articles tagged 'tutorial' on DEV Community.</description>
    <link>https://dev.to/t/tutorial</link>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tag/tutorial"/>
    <language>en</language>
    <item>
      <title>Construindo um Índice Invertido em Elixir: do zero ao TF-IDF</title>
      <dc:creator>Matheus de Camargo Marques</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:35:30 +0000</pubDate>
      <link>https://dev.to/matheuscamarques/construindo-um-indice-invertido-em-elixir-do-zero-ao-tf-idf-2eel</link>
      <guid>https://dev.to/matheuscamarques/construindo-um-indice-invertido-em-elixir-do-zero-ao-tf-idf-2eel</guid>
      <description>&lt;p&gt;Você já parou para pensar em como o Google, o Elasticsearch ou o Lucene conseguem encontrar documentos relevantes em milissegundos, mesmo com bilhões de páginas indexadas?&lt;/p&gt;

&lt;p&gt;A resposta está em uma estrutura de dados elegante chamada &lt;strong&gt;índice invertido&lt;/strong&gt;. E a boa notícia é que você pode implementá-la do zero em Elixir em menos de 100 linhas.&lt;/p&gt;

&lt;p&gt;Neste artigo, vamos construir passo a passo um índice invertido, adicionar busca booleana (AND/OR), calcular frequência de termos (TF) e, por fim, ranquear resultados com TF-IDF.&lt;/p&gt;

&lt;h2&gt;
  
  
  O que é um índice invertido?
&lt;/h2&gt;

&lt;p&gt;Se você pensar em um livro, o &lt;strong&gt;índice remissivo&lt;/strong&gt; no final lista cada termo importante e as páginas onde ele aparece. Um índice invertido faz exatamente isso, mas para documentos:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"o gato subiu no telhado"   → doc 1
"o cachorro latiu"          → doc 2

Índice invertido:
%{
  "gato"     =&amp;gt; [1],
  "subiu"    =&amp;gt; [1],
  "telhado"  =&amp;gt; [1],
  "cachorro" =&amp;gt; [2],
  "latiu"    =&amp;gt; [2]
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Em vez de, para cada busca, varrer todos os documentos (custo &lt;code&gt;O(n)&lt;/code&gt;), consultamos diretamente o mapa de termos — custo praticamente &lt;code&gt;O(1)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Simples, né? Vamos codar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preparando o projeto
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mix new indice_invertido
&lt;span class="nb"&gt;cd &lt;/span&gt;indice_invertido
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nada de dependências externas — só Elixir puro.&lt;/p&gt;

&lt;h2&gt;
  
  
  Passo 1: Tokenização
&lt;/h2&gt;

&lt;p&gt;Antes de indexar, precisamos transformar texto bruto em &lt;strong&gt;tokens&lt;/strong&gt; limpos: minúsculas, sem pontuação, sem stopwords. Esse é o trabalho da função &lt;code&gt;tokenizar/1&lt;/code&gt;, e o módulo abaixo documenta cada decisão.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;defmodule&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertido&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="nv"&gt;@moduledoc&lt;/span&gt; &lt;span class="sd"&gt;"""
  Módulo responsável por transformar texto bruto em tokens normalizados,
  etapa fundamental para a construção de um índice invertido.

  A função principal é `tokenizar/1`, que recebe uma string e devolve
  uma lista de tokens limpos — sem pontuação, sem maiúsculas e sem
  stopwords.

  ## O que são stopwords

  Stopwords são palavras tão comuns em um idioma que não ajudam a
  distinguir documentos entre si. Em português, palavras como "a",
  "o", "de", "para" aparecem em praticamente todo texto e, por isso,
  são descartadas durante a indexação. Mantê-las só aumentaria o
  tamanho do índice sem melhorar a qualidade da busca.

  As stopwords deste módulo são declaradas como atributo de módulo:

      @stopwords ~w(a o e de da do em no na um uma para com por que se)

  O sigil `~w(...)` (word list) é um atalho para criar uma lista de
  strings. Estas duas linhas são equivalentes:

      ~w(a o e de)
      ["a", "o", "e", "de"]

  ## Pipeline de tokenização

  A função `tokenizar/1` aplica quatro transformações em sequência,
  encadeadas com o operador pipe `|&amp;gt;`. O pipe pega o resultado da
  esquerda e passa como primeiro argumento da função da direita.
  Ou seja, `texto |&amp;gt; String.downcase()` é o mesmo que
  `String.downcase(texto)`.

  ### Etapa 1 — `String.downcase/1`

  Converte todo o texto para minúsculas, garantindo que "Gato",
  "gato" e "GATO" sejam tratados como o mesmo token.

      iex&amp;gt; String.downcase("O GATO Subiu")
      "o gato subiu"

  A função lida corretamente com caracteres acentuados:

      iex&amp;gt; String.downcase("AÇÃO")
      "ação"

  ### Etapa 2 — `String.replace/3` com o regex `~r/[^\\p{L}\\p{N}\\s]/u`

  Remove pontuação, símbolos e emojis, substituindo cada caractere
  indesejado por um espaço. O regex merece atenção:

    * `[^...]` — classe negada: "qualquer coisa que NÃO seja..."
    * `\\p{L}`  — qualquer letra Unicode (inclui `ç`, `é`, `日`)
    * `\\p{N}`  — qualquer número Unicode (inclui `0-9`)
    * `\\s`    — qualquer espaço em branco (espaço, tab, quebra de linha)
    * `/u`     — flag Unicode, indispensável para `\\p{...}` funcionar

  Em português: *casa com qualquer caractere que não seja letra,
  número ou espaço em branco*.

      iex&amp;gt; String.replace("olá, mundo!", ~r/[^\\p{L}\\p{N}\\s]/u, " ")
      "olá  mundo "

  A substituição é feita por **espaço**, e não por string vazia.
  Isso é importante: se apagássemos direto, `"fim.início"` viraria
  `"fimíncio"` (grudado). Com espaço, os termos permanecem separados.

      iex&amp;gt; String.replace("fim.início", ~r/[^\\p{L}\\p{N}\\s]/u, "")
      "fimíncio"

      iex&amp;gt; String.replace("fim.início", ~r/[^\\p{L}\\p{N}\\s]/u, " ")
      "fim início"

  Note que esta etapa pode gerar espaços duplos, o que é tratado
  na próxima.

  ### Etapa 3 — `String.split/3` com o regex `~r/\\s+/` e `trim: true`

  Divide a string em uma lista de tokens. O regex `\\s+` casa com
  um ou mais espaços em branco seguidos, então espaços duplos são
  tratados como um único separador. A opção `trim: true` descarta
  strings vazias no início e no final.

      iex&amp;gt; String.split("  olá  mundo  ", ~r/\\s+/)
      ["", "olá", "mundo", ""]

      iex&amp;gt; String.split("  olá  mundo  ", ~r/\\s+/, trim: true)
      ["olá", "mundo"]

  ### Etapa 4 — `Enum.reject/2` com o capture `&amp;amp;`

  Descarta os tokens que estão na lista de stopwords. `Enum.reject/2`
  percorre a lista e remove os elementos para os quais a função
  retorna `true` — é o oposto de `Enum.filter/2`.

      iex&amp;gt; Enum.reject([1, 2, 3, 4], fn x -&amp;gt; x &amp;gt; 2 end)
      [1, 2]

  O capture `&amp;amp;(&amp;amp;1 in @stopwords)` é uma forma abreviada de criar uma
  função anônima. `&amp;amp;1` refere-se ao primeiro argumento. Estas duas
  linhas são idênticas:

      Enum.reject(lista, &amp;amp;(&amp;amp;1 in @stopwords))
      Enum.reject(lista, fn token -&amp;gt; token in @stopwords end)

  O operador `in` testa se um item pertence a uma coleção:

      iex&amp;gt; "gato" in ["a", "o", "gato"]
      true

  ## Resultado final

      iex&amp;gt; IndiceInvertido.tokenizar("O gato subiu no telhado!")
      ["gato", "subiu", "telhado"]

      iex&amp;gt; IndiceInvertido.tokenizar("Ela comprou 3 maçãs, 2 peras e 1 banana!")
      ["ela", "comprou", "3", "maçãs", "2", "peras", "1", "banana"]

  Note que números são preservados por causa de `\\p{N}` no regex.
  """&lt;/span&gt;

  &lt;span class="nv"&gt;@stopwords&lt;/span&gt; &lt;span class="sx"&gt;~w(a o e de da do em no na um uma para com por que se)&lt;/span&gt;

  &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;texto&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="n"&gt;texto&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;downcase&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;~r/[^\p{L}\p{N}\s]/u&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;" "&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;~r/\s+/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ss"&gt;trim:&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nv"&gt;@stopwords&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repare no uso de &lt;code&gt;\p{L}&lt;/code&gt; e &lt;code&gt;\p{N}&lt;/code&gt; com a flag &lt;code&gt;u&lt;/code&gt; — isso garante que acentos como &lt;code&gt;é&lt;/code&gt; e &lt;code&gt;ã&lt;/code&gt; sejam preservados. Um detalhe importante ao trabalhar com português.&lt;/p&gt;

&lt;p&gt;Teste rápido:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"O gato subiu no telhado!"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"subiu"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Passo 2: Construindo o índice
&lt;/h2&gt;

&lt;p&gt;Agora a parte central. Recebemos um mapa &lt;code&gt;%{doc_id =&amp;gt; texto}&lt;/code&gt; e retornamos &lt;code&gt;%{termo =&amp;gt; [doc_ids]}&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;construir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documentos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;when&lt;/span&gt; &lt;span class="n"&gt;is_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documentos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;documentos&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;texto&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="n"&gt;texto&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;uniq&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Vamos destrinchar linha a linha.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;documentos |&amp;gt; Enum.flat_map(fn {doc_id, texto} -&amp;gt; ... end)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;Enum.flat_map/2&lt;/code&gt; percorre o mapa e, para cada par &lt;code&gt;{doc_id, texto}&lt;/code&gt;, aplica a função — mas &lt;strong&gt;achata&lt;/strong&gt; o resultado em uma única lista. Se cada documento gera uma lista de pares, o &lt;code&gt;flat_map&lt;/code&gt; junta tudo em uma lista só, em vez de uma lista de listas.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(%{&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"a b"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"c d"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt;   &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"a"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"b"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"c"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"d"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sem o &lt;code&gt;flat_map&lt;/code&gt; (usando &lt;code&gt;map&lt;/code&gt;), o resultado seria &lt;code&gt;[[{"a", 1}, {"b", 1}], [{"c", 2}, {"d", 2}]]&lt;/code&gt; — mais chato de processar depois.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Enum.uniq()&lt;/code&gt; e &lt;code&gt;Enum.map(fn termo -&amp;gt; {termo, doc_id} end)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;tokenizar/1&lt;/code&gt; devolve os tokens na ordem em que aparecem, com repetições. Se "gato" aparece três vezes no mesmo documento, não queremos três pares &lt;code&gt;{"gato", 1}&lt;/code&gt; no índice — queremos um só.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;uniq&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depois, &lt;code&gt;Enum.map/2&lt;/code&gt; transforma cada token em uma tupla &lt;code&gt;{termo, doc_id}&lt;/code&gt;, preparando os pares que serão agrupados.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Enum.group_by/3&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;Enum.group_by/3&lt;/code&gt; agrupa os pares pela chave (o termo) e coleta os valores (os &lt;code&gt;doc_id&lt;/code&gt;) em listas.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="s2"&gt;"telhado"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;O primeiro argumento é a função que extrai a chave. O segundo é a função que extrai o valor a ser coletado.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Map.new/2&lt;/code&gt; com &lt;code&gt;Enum.sort&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;Map.new/2&lt;/code&gt; reconstrói o mapa garantindo que os IDs de cada termo estejam ordenados. A ordenação ajuda nas interseções do próximo passo (busca AND), que funcionam melhor com listas ordenadas.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Testando
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;
  &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"O gato subiu no telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"O cachorro latiu para o gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"O telhado é de barro"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;construir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# %{&lt;/span&gt;
&lt;span class="c1"&gt;#   "barro"    =&amp;gt; [3],&lt;/span&gt;
&lt;span class="c1"&gt;#   "cachorro" =&amp;gt; [2],&lt;/span&gt;
&lt;span class="c1"&gt;#   "gato"     =&amp;gt; [1, 2],&lt;/span&gt;
&lt;span class="c1"&gt;#   "latiu"    =&amp;gt; [2],&lt;/span&gt;
&lt;span class="c1"&gt;#   "subiu"    =&amp;gt; [1],&lt;/span&gt;
&lt;span class="c1"&gt;#   "telhado"  =&amp;gt; [1, 3]&lt;/span&gt;
&lt;span class="c1"&gt;# }&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Passo 3: Busca booleana (AND / OR)
&lt;/h2&gt;

&lt;p&gt;Um índice sem busca é só um mapa bonito. Vamos consultá-lo.&lt;/p&gt;

&lt;h3&gt;
  
  
  Busca OR
&lt;/h3&gt;

&lt;p&gt;Encontra documentos que contenham &lt;strong&gt;qualquer&lt;/strong&gt; termo da consulta:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;buscar_or&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;consulta&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;consulta&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]))&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;uniq&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aqui reutilizamos &lt;code&gt;tokenizar/1&lt;/code&gt; na consulta — o mesmo tratamento aplicado aos documentos é aplicado à busca. Isso garante que "GATO" na query vire &lt;code&gt;"gato"&lt;/code&gt; e case com o índice.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Enum.flat_map(&amp;amp;Map.get(indice, &amp;amp;1, []))&lt;/code&gt; busca cada termo no índice. O &lt;code&gt;[]&lt;/code&gt; como padrão evita que um termo desconhecido quebre a busca — se a palavra não existe no índice, ela contribui com uma lista vazia.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="s2"&gt;"barro"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"barro"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]))&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"inexistente"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]))&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depois, &lt;code&gt;Enum.uniq/1&lt;/code&gt; remove duplicatas (um documento que contém os dois termos só aparece uma vez) e &lt;code&gt;Enum.sort/1&lt;/code&gt; mantém a ordem consistente.&lt;/p&gt;

&lt;h3&gt;
  
  
  Busca AND
&lt;/h3&gt;

&lt;p&gt;Encontra documentos que contenham &lt;strong&gt;todos&lt;/strong&gt; os termos:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;buscar_and&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;consulta&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;consulta&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]))&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;intersecao&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;

&lt;span class="k"&gt;defp&lt;/span&gt; &lt;span class="n"&gt;intersecao&lt;/span&gt;&lt;span class="p"&gt;([]),&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;defp&lt;/span&gt; &lt;span class="n"&gt;intersecao&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;lista&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;resto&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resto&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lista&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;&amp;amp;2&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;&amp;amp;2&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt; &lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aqui usamos &lt;code&gt;Enum.map/2&lt;/code&gt; (não &lt;code&gt;flat_map&lt;/code&gt;), porque queremos preservar cada lista de documentos separada — vamos intersectá-las.&lt;/p&gt;

&lt;p&gt;A função &lt;code&gt;intersecao/1&lt;/code&gt; tem duas cláusulas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;intersecao([])&lt;/code&gt; — caso base: lista vazia retorna lista vazia.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;intersecao([lista | resto])&lt;/code&gt; — pega a primeira lista e a usa como acumulador inicial do &lt;code&gt;reduce&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A expressão &lt;code&gt;&amp;amp;2 -- (&amp;amp;2 -- &amp;amp;1)&lt;/code&gt; calcula a interseção de duas listas usando apenas o operador &lt;code&gt;--&lt;/code&gt; (diferença de listas):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;&amp;amp;2 -- &amp;amp;1&lt;/code&gt; remove de &lt;code&gt;&amp;amp;2&lt;/code&gt; os elementos que estão em &lt;code&gt;&amp;amp;1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;&amp;amp;2 -- (&amp;amp;2 -- &amp;amp;1)&lt;/code&gt; remove de &lt;code&gt;&amp;amp;2&lt;/code&gt; os elementos que &lt;strong&gt;não&lt;/strong&gt; estão em &lt;code&gt;&amp;amp;1&lt;/code&gt; — sobrando exatamente a interseção.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Por isso ordenamos as listas no passo anterior: a interseção por diferença de listas depende da ordem dos elementos.&lt;/p&gt;

&lt;h3&gt;
  
  
  Uso
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_and&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato telhado"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# =&amp;gt; [1]&lt;/span&gt;
&lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_or&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato barro"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# =&amp;gt; [1, 2, 3]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Passo 4: Frequência de termos e posições
&lt;/h2&gt;

&lt;p&gt;Um índice que só diz "sim/não" não ranqueia nada. Para melhorar, guardamos &lt;strong&gt;quantas vezes&lt;/strong&gt; o termo aparece em cada documento e &lt;strong&gt;onde&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;construir_com_tf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;documentos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;documentos&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;texto&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="n"&gt;texto&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tokenizar_com_posicao&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ocorrencias&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
    &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
      &lt;span class="n"&gt;ocorrencias&lt;/span&gt;
      &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lista&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
        &lt;span class="n"&gt;posicoes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lista&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="ss"&gt;tf:&lt;/span&gt; &lt;span class="n"&gt;length&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;posicoes&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="ss"&gt;posicoes:&lt;/span&gt; &lt;span class="n"&gt;posicoes&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;
      &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;

&lt;span class="k"&gt;defp&lt;/span&gt; &lt;span class="n"&gt;tokenizar_com_posicao&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;texto&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;texto&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;downcase&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;~r/[^\p{L}\p{N}\s]/u&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;" "&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;~r/\s+/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ss"&gt;trim:&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nv"&gt;@stopwords&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;with_index&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Vamos por partes.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;tokenizar_com_posicao/1&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;É a mesma &lt;code&gt;tokenizar/1&lt;/code&gt;, mas com &lt;code&gt;Enum.with_index()&lt;/code&gt; no final. Essa função transforma cada elemento da lista em uma tupla &lt;code&gt;{elemento, índice}&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"subiu"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;with_index&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"subiu"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agora cada token carrega sua &lt;strong&gt;posição&lt;/strong&gt; no documento — informação que permite busca por frase mais adiante.&lt;/p&gt;

&lt;h3&gt;
  
  
  Primeiro &lt;code&gt;flat_map&lt;/code&gt; — desmembrando
&lt;/h3&gt;

&lt;p&gt;Para cada documento, geramos uma lista de triplas &lt;code&gt;{termo, doc_id, posicao}&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"gato gato telhado"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; 
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flat_map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt;   &lt;span class="n"&gt;txt&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;String&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;with_index&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repare que "gato" aparece &lt;strong&gt;duas vezes&lt;/strong&gt; — e é isso que queremos, porque a repetição é a informação de frequência.&lt;/p&gt;

&lt;h3&gt;
  
  
  Primeiro &lt;code&gt;group_by&lt;/code&gt; — agrupando por termo
&lt;/h3&gt;

&lt;p&gt;Agrupamos todas as triplas pelo termo. Note o padrão &lt;code&gt;{termo, _, _}&lt;/code&gt; — usamos &lt;code&gt;_&lt;/code&gt; para ignorar o &lt;code&gt;doc_id&lt;/code&gt; e a posição, já que estamos agrupando apenas pelo termo.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;%{&lt;/span&gt;
  &lt;span class="s2"&gt;"gato"&lt;/span&gt;    &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
  &lt;span class="s2"&gt;"telhado"&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Segundo &lt;code&gt;group_by&lt;/code&gt; dentro do &lt;code&gt;Map.new&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Para cada termo, precisamos agrupar as ocorrências por documento. O padrão &lt;code&gt;{_, doc_id, _}&lt;/code&gt; ignora o termo e a posição.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="o"&gt;...&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;group_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;%{&lt;/span&gt;
  &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
  &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agora, dentro do &lt;code&gt;Map.new/2&lt;/code&gt; mais interno, extraímos apenas as posições e calculamos o &lt;code&gt;tf&lt;/code&gt; (quantas ocorrências o termo teve naquele documento):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;posicoes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lista&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="ss"&gt;tf:&lt;/span&gt; &lt;span class="n"&gt;length&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;posicoes&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="ss"&gt;posicoes:&lt;/span&gt; &lt;span class="n"&gt;posicoes&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;O &lt;code&gt;length(posicoes)&lt;/code&gt; é o &lt;strong&gt;term frequency&lt;/strong&gt; para aquele par termo/documento. O &lt;code&gt;posicoes&lt;/code&gt; guarda onde o termo apareceu, permitindo busca por frase depois.&lt;/p&gt;

&lt;h3&gt;
  
  
  Resultado
&lt;/h3&gt;

&lt;p&gt;Agora o índice fica mais rico:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;%{&lt;/span&gt;
  &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="ss"&gt;tf:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ss"&gt;posicoes:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
  &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="ss"&gt;tf:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ss"&gt;posicoes:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Com as &lt;strong&gt;posições&lt;/strong&gt;, você pode até implementar busca por frase — basta verificar se as posições dos termos são adjacentes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Passo 5: Ranking com TF-IDF
&lt;/h2&gt;

&lt;p&gt;Agora o pulo do gato (trocadilho inevitável). &lt;strong&gt;TF-IDF&lt;/strong&gt; é uma métrica clássica que combina:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;TF&lt;/strong&gt; (Term Frequency): quanto mais o termo aparece no documento, mais relevante ele é para aquele documento.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IDF&lt;/strong&gt; (Inverse Document Frequency): quanto mais raro o termo no corpus inteiro, mais informativo ele é.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Palavras como "de" aparecem em tudo — IDF baixo. Palavras como "criptografia" aparecem em poucos documentos — IDF alto.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="n"&gt;buscar_ranqueado&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;total_docs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;consulta&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="n"&gt;termos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tokenizar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;consulta&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

  &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;termos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;%{},&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;acc&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
      &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;termo&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
        &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;acc&lt;/span&gt;
        &lt;span class="n"&gt;docs&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
          &lt;span class="n"&gt;idf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="ss"&gt;:math&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;total_docs&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;map_size&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
          &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;acc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;&lt;span class="ss"&gt;tf:&lt;/span&gt; &lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;
            &lt;span class="no"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tf&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;idf&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;&amp;amp;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;tf&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;idf&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
          &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;end&lt;/span&gt;
    &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

  &lt;span class="n"&gt;scores&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;Float&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;code&gt;Enum.reduce(termos, %{}, fn termo, acc -&amp;gt; ... end)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Percorremos cada termo da consulta, acumulando um mapa &lt;code&gt;%{doc_id =&amp;gt; score}&lt;/code&gt;. O acumulador começa vazio e vai sendo atualizado termo a termo.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;case Map.get(indice, termo) do ... end&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Se o termo não existe no índice, &lt;code&gt;Map.get&lt;/code&gt; retorna &lt;code&gt;nil&lt;/code&gt; e simplesmente mantemos o acumulador. Se existe, calculamos a contribuição dele.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;idf = :math.log(total_docs / map_size(docs))&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;O IDF é o logaritmo natural da razão entre o total de documentos e o número de documentos que contêm o termo. &lt;code&gt;map_size(docs)&lt;/code&gt; conta quantos documentos (chaves) o termo tem.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="ss"&gt;:math&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mf"&gt;0.4054651081081644&lt;/span&gt;

&lt;span class="n"&gt;iex&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="ss"&gt;:math&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="mf"&gt;1.0986122886681098&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Um termo que aparece em 1 de 3 documentos tem IDF maior que um que aparece em 2 de 3 — confirmando que termos raros são mais informativos.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Enum.reduce(docs, acc, fn {doc_id, %{tf: tf}}, a -&amp;gt; ... end)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Para cada documento onde o termo aparece, desestruturamos o padrão &lt;code&gt;%{tf: tf}&lt;/code&gt; — pattern matching direto na cabeça do &lt;code&gt;fn&lt;/code&gt;, extraindo o &lt;code&gt;tf&lt;/code&gt; do mapa.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;Map.update(a, doc_id, tf * idf, &amp;amp;(&amp;amp;1 + tf * idf))&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;Map.update/4&lt;/code&gt; atualiza o score do documento:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Se o &lt;code&gt;doc_id&lt;/code&gt; ainda não está no mapa, insere com valor &lt;code&gt;tf * idf&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Se já está (o documento tem mais de um termo da consulta), soma &lt;code&gt;tf * idf&lt;/code&gt; ao valor existente com a função &lt;code&gt;&amp;amp;(&amp;amp;1 + tf * idf)&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;É por isso que documentos que contêm &lt;strong&gt;todos&lt;/strong&gt; os termos da consulta acumulam score maior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ordenação
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="o"&gt;|&amp;gt;&lt;/span&gt; &lt;span class="no"&gt;Enum&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sort_by&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;_doc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Enum.sort_by/2&lt;/code&gt; ordena pelos scores, mas com &lt;code&gt;-score&lt;/code&gt; invertemos — do maior para o menor. Sem isso, seria ordem crescente.&lt;/p&gt;

&lt;h3&gt;
  
  
  Uso
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;construir_com_tf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_ranqueado&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato telhado"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# =&amp;gt; [{1, 0.811}, {2, 0.405}]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Documentos que contêm &lt;strong&gt;os dois&lt;/strong&gt; termos pontuam mais alto — exatamente o que esperamos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Passo 6: Testes
&lt;/h2&gt;

&lt;p&gt;Nenhum tutorial sério termina sem testes. Com ExUnit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight elixir"&gt;&lt;code&gt;&lt;span class="k"&gt;defmodule&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertidoTest&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
  &lt;span class="kn"&gt;use&lt;/span&gt; &lt;span class="no"&gt;ExUnit&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="no"&gt;Case&lt;/span&gt;
  &lt;span class="n"&gt;alias&lt;/span&gt; &lt;span class="no"&gt;IndiceInvertido&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ss"&gt;as:&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;

  &lt;span class="nv"&gt;@docs&lt;/span&gt; &lt;span class="p"&gt;%{&lt;/span&gt;
    &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"O gato subiu no telhado"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"O cachorro latiu para o gato"&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="n"&gt;test&lt;/span&gt; &lt;span class="s2"&gt;"constrói índice básico"&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;construir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;@docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gato"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"cachorro"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;

  &lt;span class="n"&gt;test&lt;/span&gt; &lt;span class="s2"&gt;"busca AND"&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;construir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;@docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;assert&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_and&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato cachorro"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;assert&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_and&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"gato telhado"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;

  &lt;span class="n"&gt;test&lt;/span&gt; &lt;span class="s2"&gt;"busca OR"&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="n"&gt;indice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;construir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;@docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;assert&lt;/span&gt; &lt;span class="no"&gt;II&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;buscar_or&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"cachorro telhado"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;O &lt;code&gt;alias IndiceInvertido, as: II&lt;/code&gt; permite referenciar o módulo como &lt;code&gt;II&lt;/code&gt; nos testes, deixando o código mais curto. O &lt;code&gt;@docs&lt;/code&gt; é um atributo de módulo — uma constante compartilhada entre os testes.&lt;/p&gt;

&lt;p&gt;Rode com &lt;code&gt;mix test&lt;/code&gt; e pronto.&lt;/p&gt;

&lt;h2&gt;
  
  
  Onde ir a partir daqui
&lt;/h2&gt;

&lt;p&gt;O que construímos é a fundação. Alguns caminhos naturais para evoluir:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stemming em português&lt;/strong&gt; — reduzir &lt;code&gt;gatos&lt;/code&gt; → &lt;code&gt;gato&lt;/code&gt; com a lib &lt;a href="https://hex.pm/packages/stemmer" rel="noopener noreferrer"&gt;&lt;code&gt;stemmer&lt;/code&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GenServer&lt;/strong&gt; — encapsular o índice em um processo para consultas concorrentes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Task.async_stream&lt;/code&gt;&lt;/strong&gt; — paralelizar a indexação de milhares de documentos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistência&lt;/strong&gt; — usar &lt;code&gt;:dets&lt;/code&gt;, &lt;code&gt;:ets&lt;/code&gt; ou &lt;code&gt;:mnesia&lt;/code&gt; para não reconstruir o índice a cada boot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Busca por frase&lt;/strong&gt; — usando as posições que já guardamos.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusão
&lt;/h2&gt;

&lt;p&gt;Índice invertido é uma daquelas estruturas que parecem mágica até você escrever a sua. Em Elixir, o poder do &lt;code&gt;Enum&lt;/code&gt;, &lt;code&gt;Map&lt;/code&gt; e pattern matching torna a implementação surpreendentemente enxuta e legível.&lt;/p&gt;

&lt;p&gt;O código completo está no &lt;a href="https://github.com/matheuscamarques/my_research" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; — fique à vontade para fazer fork e experimentar.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;E você?&lt;/strong&gt; Já usou índices invertidos em algum projeto? Conta aqui nos comentários o que achou — e se quiser, no próximo artigo posso mostrar como transformar isso num servidor GenServer concorrente. &lt;/p&gt;

</description>
      <category>elixir</category>
      <category>tutorial</category>
      <category>search</category>
      <category>programming</category>
    </item>
    <item>
      <title>Cloudflare libera Clef y Clef-flash bajo licencia Apache 2.0</title>
      <dc:creator>lu1tr0n</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:31:16 +0000</pubDate>
      <link>https://dev.to/lu1tr0n/cloudflare-libera-clef-y-clef-flash-bajo-licencia-apache-20-4nmn</link>
      <guid>https://dev.to/lu1tr0n/cloudflare-libera-clef-y-clef-flash-bajo-licencia-apache-20-4nmn</guid>
      <description>&lt;p&gt;Cloudflare acaba de lanzar dos modelos de inteligencia artificial que no generan texto: deciden. Clef y Clef-flash clasifican, enrutan y escalan tareas con probabilidades calculadas, y la empresa asegura que ya superan al modelo que detonó la moda de los &lt;strong&gt;modelos de decisión&lt;/strong&gt; en las últimas semanas.&lt;/p&gt;

&lt;p&gt;Los publicó el 2 de octubre de 2026 bajo licencia Apache 2.0 en Hugging Face, y los dejó disponibles de inmediato en Workers AI, su plataforma de inferencia. Además estrenó una plataforma de ajuste fino con aprendizaje por refuerzo para que cualquier equipo adapte Clef a su propio flujo de trabajo.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Cloudflare lanzó Clef y Clef-flash, modelos de decisión de código abierto bajo Apache 2.0, disponibles en Workers AI.- Clef lidera el Jev Decision Index y vence a Jev en 3 de 4 flujos de trabajo evaluados por Typesafe AI.- A diferencia de Jev, Clef procesa imágenes gracias a un encoder de visión y tiene contexto de 64k tokens frente a 32k.- Clasificar un dominio tardó 2,2s con Clef y 4,7s con el LLM gpt-oss-120b, con menos categorías devueltas.- Suma además una plataforma de ajuste fino con aprendizaje por refuerzo para personalizar Clef en casos propios.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Qué es Clef
&lt;/h2&gt;

&lt;p&gt;Clef es uno de los nuevos modelos de decisión que Cloudflare entrena y publica en código abierto para producir salidas estructuradas con probabilidades, como clasificar un mensaje, una imagen o un dominio web, sin reentrenarlo cada vez que aparece una categoría nueva. Corre en Workers AI y es compatible con la API de Jev.&lt;br&gt;
Clef combina texto e imagen; Jev, su referencia directa, solo procesa texto.&lt;br&gt;
El nombre viene de la música. Una clave (clef, en inglés) es el símbolo que fija qué nota representa cada línea del pentagrama, y Cloudflare usa la misma idea para su modelo: Clef fija el dominio de la decisión y las acciones que siguen después, y de paso las dos primeras letras guiñan a la propia compañía.&lt;/p&gt;
&lt;h2&gt;
  
  
  Qué pasó
&lt;/h2&gt;

&lt;p&gt;Cloudflare detalló el lanzamiento en su &lt;a href="https://blog.cloudflare.com/clef-decision-models/" rel="noopener noreferrer"&gt;blog oficial&lt;/a&gt;, donde publicó los pesos de Clef y Clef-flash en Hugging Face bajo la &lt;a href="https://www.apache.org/licenses/LICENSE-2.0" rel="noopener noreferrer"&gt;licencia Apache 2.0&lt;/a&gt;, la misma que usan la mayoría de proyectos de infraestructura abierta porque permite uso comercial sin pedir permiso. Ese mismo día, ambos modelos quedaron disponibles en &lt;a href="https://developers.cloudflare.com/workers-ai/" rel="noopener noreferrer"&gt;Workers AI&lt;/a&gt;, la plataforma de inferencia de Cloudflare, para que cualquier desarrollador los pruebe sin desplegar infraestructura propia.&lt;/p&gt;

&lt;p&gt;La compañía también anunció una plataforma de ajuste fino con aprendizaje por refuerzo (RL) para Clef, pensada para que los equipos lo entrenen sobre sus propios datos sin empezar desde cero. El anuncio llegó en medio del interés que generó semanas atrás el Jev System One, el modelo de decisión de Typesafe AI que popularizó el término fuera de los papers académicos.&lt;/p&gt;
&lt;h2&gt;
  
  
  El auge reciente de los modelos de decisión
&lt;/h2&gt;

&lt;p&gt;Los clasificadores binarios existen hace años, pero el término "modelos de decisión" se extendió apenas en las últimas semanas, cuando Typesafe AI mostró que Jev System One devolvía respuestas tipadas con probabilidades para decisiones puntuales, en lugar de texto libre. La diferencia con un LLM de uso general es clara: un modelo de decisión responde dentro de un conjunto fijo de categorías, con un número que indica qué tan seguro está, sin necesidad de reentrenarse si aparece una categoría nueva dentro de ese mismo conjunto. Este tipo de modelos de clasificación gana terreno porque es más barato de operar que un LLM general para tareas repetitivas.&lt;/p&gt;

&lt;p&gt;Un ejemplo simple lo explica mejor que la definición. Si un equipo de soporte recibe un mensaje, puede pasarlo por un modelo de decisión y preguntar si es urgente y qué equipo debería atenderlo. El modelo devuelve una respuesta tipada con probabilidades (urgente: 92%, equipo de facturación: 78%) que el código puede usar para enrutar el ticket, disparar una alerta o esperar a que decida una persona. Ya no hace falta que un humano esté en el medio de cada decisión agéntica.&lt;/p&gt;
&lt;h2&gt;
  
  
  Cómo compite Clef en precisión y latencia
&lt;/h2&gt;

&lt;p&gt;Clef tiene dos diferencias estructurales frente a Jev. La primera es un encoder de visión, que le permite clasificar imágenes además de texto, algo que Jev todavía no hace. La segunda es el contexto, de 64.000 tokens frente a los 32.000 de Jev, lo que permite meter más información de entrada antes de pedir una clasificación.&lt;/p&gt;

&lt;p&gt;Cloudflare publicó resultados de varios benchmarks usados para medir decisiones automatizadas, recogidos en el &lt;a href="https://blog.cloudflare.com/clef-decision-models/" rel="noopener noreferrer"&gt;anuncio oficial&lt;/a&gt;. Estos son algunos de los más representativos:&lt;br&gt;
BenchmarkClefClef-flashJevBFCL (case exact)98.4798.7695.75API-Bank (accuracy)91.9393.1188.19BANKING77 (macro-F1)94.2090.9379.74CLINC150+OOS (macro-F1)97.4366.7789.27When2Call (accuracy)72.3765.5880.97PhishNChips (accuracy)79.6075.0562.55&lt;br&gt;
Clef gana en cuatro de las seis pruebas de la tabla. When2Call es la excepción más notoria, ahí Jev saca 80.97 contra el 72.37 de Clef, lo que confirma que ningún modelo domina todos los escenarios por igual.&lt;/p&gt;

&lt;p&gt;La latencia es donde más se nota la diferencia de arquitectura. Workers AI corre los modelos sobre la infraestructura de borde de Cloudflare, lo que recorta el tiempo de ida y vuelta frente a un servicio centralizado en una sola región.&lt;br&gt;
ModeloLatencia medianaLatencia p95Clef209.3 ms238.6 msClef-flash38.8 ms122.4 msJev524.1 ms536.0 msClef-flash respondió en 38,8 ms de mediana frente a 524,1 ms de Jev, según Cloudflare.&lt;br&gt;
&lt;a href="https://blog.cloudflare.com/clef-decision-models/" rel="noopener noreferrer"&gt;Cloudflare también probó&lt;/a&gt; los modelos en flujos de trabajo completos, no solo en benchmarks aislados. En procesamiento de facturas, Clef anotó 64.7 puntos contra 61.8 de Jev; en servicio al cliente, Clef-flash lideró con 77 frente a 76.3 de Clef y 76.0 de Jev. La única categoría donde Jev ganó con claridad fue la observabilidad de trazas de agentes, con 71.6 puntos frente a los 68.5 de Clef y 69.8 de Clef-flash.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;💭 Clave:&lt;/strong&gt; Clef no gana en todo. Jev sigue por delante en When2Call, en el benchmark BRIGHT y en la observabilidad de trazas de agentes, así que la elección depende de qué parte del flujo de trabajo importa más en cada caso.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;El caso real que Cloudflare usa de ejemplo es el de su propio equipo de inteligencia de amenazas. Al pasarle un dominio a Clef, con la herramienta Browser Run para renderizarlo primero, el modelo devuelve categorías con probabilidad: una web puede salir con 95% de probabilidad de ser de moda, 85% de comercio electrónico y menos de 1% de phishing. Clasificar ese dominio completo, buscar la página, renderizarla y clasificarla, &lt;a href="https://blog.cloudflare.com/clef-decision-models/" rel="noopener noreferrer"&gt;tomó 2,2 segundos con Clef&lt;/a&gt;. El mismo flujo con &lt;strong&gt;gpt-oss-120b&lt;/strong&gt;, un LLM de uso general, tardó 4,7 segundos y devolvió solo dos categorías en vez de las múltiples que entrega Clef.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight dot"&gt;&lt;code&gt;&lt;span class="nv"&gt;flowchart&lt;/span&gt; &lt;span class="nv"&gt;TD&lt;/span&gt;
    &lt;span class="nv"&gt;A&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Entrada: texto o imagen"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;B&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Clef en Workers AI"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;B&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;C&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Salida estructurada con probabilidades"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;C&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nv"&gt;D&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"Confianza alta?"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nv"&gt;D&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="s2"&gt;"Si"&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;E&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Accion automatica"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
    &lt;span class="nv"&gt;D&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;|&lt;/span&gt;&lt;span class="s2"&gt;"No"&lt;/span&gt;&lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="nv"&gt;F&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Escala a un humano"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;El diagrama resume el circuito completo. La entrada llega a Clef, el modelo devuelve una salida tipada con probabilidades, y el código decide si actúa solo o pasa el caso a una persona según el umbral de confianza que defina cada equipo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Impacto y análisis
&lt;/h2&gt;

&lt;p&gt;La compatibilidad con la API de Jev es la decisión más práctica del lanzamiento. Un equipo que ya integró Jev System One puede apuntar sus llamadas a Clef sin reescribir la lógica de enrutamiento, lo que convierte a Cloudflare en una alternativa de bajo costo de cambio más que en una ruptura técnica.&lt;/p&gt;

&lt;p&gt;El soporte de imágenes abre casos de uso que Jev no cubre hoy, como moderación de contenido visual, clasificación de documentos escaneados o revisión de capturas de pantalla dentro de un flujo de soporte. Sumado al contexto de 64.000 tokens, Clef puede recibir más contexto por decisión sin fragmentar la entrada en varias llamadas.&lt;/p&gt;

&lt;p&gt;La plataforma de ajuste fino con aprendizaje por refuerzo es, además, el primer paso de Cloudflare hacia cobrar por personalización y no solo por inferencia. Clef y Clef-flash son gratis de descargar y correr donde quiera el usuario, gracias a la licencia Apache 2.0, pero entrenarlos sobre datos propios dentro de Workers AI es un producto separado.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qué sigue
&lt;/h2&gt;

&lt;p&gt;Cloudflare no publicó una fecha para abrir la plataforma de RL a todos los clientes de Workers AI, aunque el anuncio ya la describe como disponible junto con los modelos. Lo que sí confirmó es que su propio equipo de inteligencia de amenazas sigue usando Clef en producción para clasificar dominios, y que seguirá publicando resultados contra la Jev Decision Index a medida que entrenen nuevas versiones.&lt;/p&gt;

&lt;p&gt;Al quedar los pesos abiertos en Hugging Face, la comunidad puede ahora comparar a Clef contra otros modelos de decisión que ya aparecen en el mismo benchmark, como DiffusionGemma, Jev Kev 9B y Laya, sin depender de que Cloudflare publique el número.&lt;/p&gt;

&lt;p&gt;📖 Resumen en Telegram: &lt;a href="https://telegra.ph/Cloudflare-libera-Clef-y-Clef-flash-bajo-licencia-Apache-20-10-02" rel="noopener noreferrer"&gt;Ver resumen&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Probalo vos: bajá los pesos de Clef desde Hugging Face o probalo directo en Workers AI con tu propia cuenta de Cloudflare para ver cuánto tarda en tu caso real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preguntas frecuentes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ¿Qué son los modelos de decisión?
&lt;/h3&gt;

&lt;p&gt;Son modelos de IA que devuelven respuestas tipadas con probabilidades dentro de un conjunto fijo de categorías, en lugar de generar texto libre. Se usan para automatizar decisiones repetitivas como enrutar un ticket o clasificar un dominio.&lt;/p&gt;

&lt;h3&gt;
  
  
  ¿Clef reemplaza a un LLM como GPT o Claude?
&lt;/h3&gt;

&lt;p&gt;No, cubre una tarea distinta. Un LLM razona en lenguaje abierto y puede generar texto o llamar herramientas; Clef solo clasifica dentro de categorías conocidas, más rápido y con un costo menor por llamada.&lt;/p&gt;

&lt;h3&gt;
  
  
  ¿Clef y Clef-flash son gratis?
&lt;/h3&gt;

&lt;p&gt;Los pesos están publicados en Hugging Face bajo licencia Apache 2.0, así que correrlos localmente no tiene costo de licencia. Usarlos hospedados en Workers AI sigue el esquema de precios de esa plataforma.&lt;/p&gt;

&lt;h3&gt;
  
  
  ¿Qué diferencia hay entre Clef y Clef-flash?
&lt;/h3&gt;

&lt;p&gt;Clef-flash es la versión optimizada para latencia: respondió en 38,8 ms de mediana frente a los 209,3 ms de Clef en las pruebas publicadas por Cloudflare, aunque pierde precisión en algunos benchmarks como CLINC150+OOS.&lt;/p&gt;

&lt;h3&gt;
  
  
  ¿Qué es la Jev Decision Index?
&lt;/h3&gt;

&lt;p&gt;Es el benchmark público que compara modelos de decisión, entre ellos Jev, Clef, Clef-flash, DiffusionGemma, Jev Kev 9B y Laya, usando las mismas pruebas de clasificación y flujos de trabajo.&lt;/p&gt;

&lt;h3&gt;
  
  
  ¿Puedo entrenar mi propia versión de Clef?
&lt;/h3&gt;

&lt;p&gt;Sí, Cloudflare lanzó junto con los modelos una plataforma de ajuste fino con aprendizaje por refuerzo pensada para adaptar Clef a categorías y datos propios de cada equipo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Referencias
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.cloudflare.com/clef-decision-models/" rel="noopener noreferrer"&gt;Cloudflare Blog&lt;/a&gt;: anuncio oficial de Clef, Clef-flash y la plataforma de ajuste fino con RL.- &lt;a href="https://www.apache.org/licenses/LICENSE-2.0" rel="noopener noreferrer"&gt;Apache License 2.0&lt;/a&gt;: texto de la licencia bajo la que Cloudflare publicó los pesos de ambos modelos.- &lt;a href="https://developers.cloudflare.com/workers-ai/" rel="noopener noreferrer"&gt;Cloudflare Workers AI&lt;/a&gt;: documentación de la plataforma de inferencia donde corren Clef y Clef-flash.- &lt;a href="https://huggingface.co" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt;: repositorio donde Cloudflare publicó los pesos abiertos de Clef y Clef-flash.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;📱 &lt;strong&gt;¿Te gusta este contenido?&lt;/strong&gt; Únete a nuestro canal de Telegram &lt;a href="https://t.me/programacion" rel="noopener noreferrer"&gt;@programacion&lt;/a&gt; donde publicamos a diario lo más relevante de tecnología, IA y desarrollo. Resúmenes rápidos, contenido fresco todos los días.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Swap tokens from your agent wallet: quote first, then execute</title>
      <dc:creator>OpenClaw Cash</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:08:55 +0000</pubDate>
      <link>https://dev.to/openclawcash/swap-tokens-from-your-agent-wallet-quote-first-then-execute-2ema</link>
      <guid>https://dev.to/openclawcash/swap-tokens-from-your-agent-wallet-quote-first-then-execute-2ema</guid>
      <description>&lt;p&gt;An agent that holds SOL and owes USDC is stuck: it can send what it holds, but the invoice is in a token it does not hold. Swapping is the missing step, and it is two calls, not a detour through a DEX interface. This is the whole flow over the agent API, with every route and field copied out of the public reference at &lt;a href="https://openclawcash.com/docs" rel="noopener noreferrer"&gt;https://openclawcash.com/docs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Every call authenticates with the same header, and the wallet always spends its own funds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;X-Agent-Key: occ_your_api_key
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  1. Get the wallet id
&lt;/h2&gt;

&lt;p&gt;Everything else selects a wallet, so start where the ids are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://openclawcash.com/api/agent/wallets?includeBalances=true"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response lists each wallet with &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;label&lt;/code&gt;, &lt;code&gt;address&lt;/code&gt;, &lt;code&gt;network&lt;/code&gt; and &lt;code&gt;chain&lt;/code&gt;, plus the native balance when you ask for it. Keep the &lt;code&gt;id&lt;/code&gt;, it is what the later calls take.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Quote before anything moves
&lt;/h2&gt;

&lt;p&gt;The quote route prices the swap and signs nothing. It needs &lt;code&gt;X-Agent-Key&lt;/code&gt; like every other agent call, and it is the safe place to look before you commit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://openclawcash.com/api/agent/quote?network=solana-mainnet"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "chain": "solana",
    "walletId": "YUZE66Z",
    "tokenIn": "SOL",
    "tokenOut": "USDC",
    "amountIn": "10000000"
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The answer carries &lt;code&gt;amountOut&lt;/code&gt;, &lt;code&gt;amountOutHuman&lt;/code&gt;, &lt;code&gt;amountIn&lt;/code&gt;, &lt;code&gt;amountInHuman&lt;/code&gt;, &lt;code&gt;route&lt;/code&gt;, &lt;code&gt;feePercent&lt;/code&gt;, &lt;code&gt;dex&lt;/code&gt; and &lt;code&gt;network&lt;/code&gt;. No transaction is built, no fee is taken and nothing is signed.&lt;/p&gt;

&lt;p&gt;On EVM the same route takes the chain as a query parameter instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://openclawcash.com/api/agent/quote?network=polygon-mainnet"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "walletId": "Q7X2K9P",
    "tokenIn": "WETH",
    "tokenOut": "USDC",
    "amountIn": "100000000000000000"
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;amountIn&lt;/code&gt; is in base units in both cases: lamports on Solana, so &lt;code&gt;10000000&lt;/code&gt; is 0.01 SOL, and wei on EVM, so &lt;code&gt;1000000000000000000&lt;/code&gt; is 1 WETH. That is easy to get wrong, which is why the quote hands &lt;code&gt;amountInHuman&lt;/code&gt; and &lt;code&gt;amountOutHuman&lt;/code&gt; back to you.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Read the policy that will check the swap
&lt;/h2&gt;

&lt;p&gt;Policies are checked before anything is signed, so read them before you are refused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://openclawcash.com/api/agent/policies &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each entry pairs a wallet with its policies and their current usage. A &lt;code&gt;daily_spending_limit&lt;/code&gt; comes back with &lt;code&gt;config.amount&lt;/code&gt; and a &lt;code&gt;usage&lt;/code&gt; block holding &lt;code&gt;spent&lt;/code&gt; and &lt;code&gt;limit&lt;/code&gt; for the window, so you can see how much of the cap the swap will consume before you send it.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Execute
&lt;/h2&gt;

&lt;p&gt;Slippage lives here and not on the quote. This is the call that moves funds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://openclawcash.com/api/agent/swap &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "chain": "solana",
    "walletId": "YUZE66Z",
    "tokenIn": "SOL",
    "tokenOut": "USDC",
    "amountIn": "10000000",
    "slippage": 0.5
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response gives you &lt;code&gt;txHash&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt;, &lt;code&gt;dex&lt;/code&gt;, &lt;code&gt;amountOut&lt;/code&gt;, &lt;code&gt;amountOutMin&lt;/code&gt;, &lt;code&gt;fee&lt;/code&gt; and &lt;code&gt;feePercent&lt;/code&gt;. &lt;code&gt;amountOutMin&lt;/code&gt; is the guard that came out of the slippage value you sent.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Confirm what landed
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://openclawcash.com/api/agent/transactions?walletId=YUZE66Z&amp;amp;chain=solana"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"X-Agent-Key: occ_your_api_key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;History merges on-chain and app-recorded rows, and each row carries its own &lt;code&gt;network&lt;/code&gt;, so an EVM wallet can scope to one chain with &lt;code&gt;network=base-mainnet&lt;/code&gt; or merge the bucket with &lt;code&gt;network=all&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to watch
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Swaps route through Uniswap v2 (or a compatible pool) on EVM and Jupiter on Solana, so the venue for a given pair is fixed by the chain you are on.&lt;/li&gt;
&lt;li&gt;Spending an ERC-20 on EVM can need &lt;code&gt;POST /api/agent/approve&lt;/code&gt; first, which approves a spender contract for that token. The route is EVM only and it takes &lt;code&gt;tokenAddress&lt;/code&gt;, &lt;code&gt;spender&lt;/code&gt; and &lt;code&gt;amount&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The quote is a price, not an action. Nothing moves until you call the swap, and the price can move between the two calls, which is what the slippage value and &lt;code&gt;amountOutMin&lt;/code&gt; are there for.&lt;/li&gt;
&lt;li&gt;Any valid ERC-20 contract or SPL mint works in wallet operations, not only the tokens the supported-tokens route recommends.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>crypto</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Reimplementing a Trie Reminded Me How Autocomplete Actually Works</title>
      <dc:creator>Darshan Turakhia</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:06:26 +0000</pubDate>
      <link>https://dev.to/darshan_turakhia/reimplementing-a-trie-reminded-me-how-autocomplete-actually-works-5ghd</link>
      <guid>https://dev.to/darshan_turakhia/reimplementing-a-trie-reminded-me-how-autocomplete-actually-works-5ghd</guid>
      <description>&lt;p&gt;I spent an embarrassing amount of time last week reimplementing something I'd "known" since a data structures class a decade ago: the trie. Typing it out from scratch again reminded me why it's one of the few data structures that actually earns the "elegant" label people throw around too easily.&lt;/p&gt;

&lt;p&gt;Here's the pitch in one sentence: a trie stores strings by character, along shared paths from a root, so two words with the same prefix walk the exact same nodes until they stop agreeing. Insert "car," "card," and "care," and you get one shared path for "c-a-r," which then splits into three different endings. Nothing about adding "care" touches the nodes "car" and "card" already built.&lt;/p&gt;

&lt;p&gt;That sharing is the whole point, and it's easy to undersell how much it buys you. A hash set can tell you "yes, this exact string is in the collection." That's it. It cannot tell you "give me everything that starts with these three letters" without scanning the entire set and checking each entry one by one. A trie answers that second question by walking to the end of the prefix and then just... looking at what's underneath. Lookup cost depends on how long your prefix is, not on how many million other strings happen to be sitting next to it in memory. That's the entire mechanism behind every autocomplete dropdown you've ever typed into: walk the prefix, then collect every real word hanging off wherever you land.&lt;/p&gt;

&lt;p&gt;The part that actually bit me while rebuilding this: you need an explicit marker for "a real word ends here." It sounds obvious written down, but it's the single easiest thing to forget, and the bug it causes is sneaky. Say you only ever insert "card." The path c → a → r → d exists in your trie. All four nodes are real. But "car" was never inserted as its own entry, it's just a prefix that happens to lead somewhere. If your node doesn't carry a boolean flag for "an insertion actually terminated here," you have no way to distinguish a real stored word from a string that merely happens to be a prefix of a longer one. I've seen (and once written) autocomplete bugs where a partial prefix gets suggested as a complete match because of exactly this missing flag.&lt;/p&gt;

&lt;p&gt;The tradeoff nobody mentions in the five-minute version of this concept: a naive implementation wastes a lot of memory. The textbook approach gives every node a fixed-size array, one slot per possible next character. Twenty-six slots if you're doing lowercase English, more if you need digits or punctuation or, god forbid, full Unicode. Most of those slots sit empty on any node that doesn't branch in every direction, and that waste adds up fast across a trie with real-world vocabulary in it. The fix is one of two things: back each node with a hash map instead of a fixed array, so you only pay for children that actually exist, or compress the whole structure with something like a radix trie, which collapses long runs of single-child nodes into one node holding a whole substring instead of one character. "Cardboard" doesn't need five separate one-character hops after "card," it needs one node labeled "board."&lt;/p&gt;

&lt;p&gt;Autocomplete is the use case everyone reaches for first, and it's a good one, but it's not the only place this shows up. Spell-checkers use the same walk to confirm a word exists in a dictionary and to find near-miss suggestions by exploring nearby paths. IP routers use a binary version of the same idea, walking a trie over address bits instead of letters, to do longest-prefix-match routing: find the most specific rule that matches a destination address. Different alphabet, same shared-prefix-sharing-a-path trick underneath.&lt;/p&gt;

&lt;p&gt;I ended up writing both a longer explanation and a small interactive visualizer, mostly because I wanted to actually watch the branching happen rather than just trust my own mental model of it. If you want to poke at it yourself, &lt;a href="https://nodique.com/guides/what-is-a-trie" rel="noopener noreferrer"&gt;the guide&lt;/a&gt; walks through the insert/search mechanics and the end-of-word gotcha in more detail, and &lt;a href="https://nodique.com/tools/trie-visualizer" rel="noopener noreferrer"&gt;the visualizer&lt;/a&gt; lets you insert your own words and watch the prefix lookup highlight in real time, suggestions and all.&lt;/p&gt;

&lt;p&gt;If you've never implemented one, it's a genuinely good weekend exercise. Small enough to build in an hour. Annoying enough, in the best way, that you'll start second-guessing every search box you touch afterward.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>computerscience</category>
      <category>webdev</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>PNG to WebP 2026: The Complete Guide to Smarter Images</title>
      <dc:creator>Sajith Madushanka</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:01:36 +0000</pubDate>
      <link>https://dev.to/sajith_madushanka_3ce9e28/png-to-webp-2026-the-complete-guide-to-smarter-images-lop</link>
      <guid>https://dev.to/sajith_madushanka_3ce9e28/png-to-webp-2026-the-complete-guide-to-smarter-images-lop</guid>
      <description>&lt;p&gt;Understanding PNG and WebP Image Formats&lt;/p&gt;

&lt;p&gt;&lt;a href="https://png2webpnow.store/" rel="noopener noreferrer"&gt;PNG and WebP&lt;/a&gt; are common formats for digital images. Each format serves a different purpose online. PNG preserves clear details and sharp graphic edges. It also supports transparent backgrounds &lt;br&gt;
for many designs. WebP focuses on efficient image delivery online. It can support transparency and strong visual quality. Its efficient compression can create smaller image files. Smaller files can simplify website image management. PNG files can become large with detailed graphics. This happens often with screenshots and illustrations. WebP offers a useful alternative for those images. PNG to WebP changes the image format efficiently. The original image purpose remains largely unchanged. Users can complete this process with simple tools. Most converters require only a few actions. First, choose the PNG image for conversion. Next, upload it into your chosen converter. Then, select WebP as the output format. Finally, save the converted image for later use. Always inspect the result before publishing it. Check colors, edges, transparency, and dimensions carefully. These checks help maintain a polished appearance. Understanding both formats makes conversion much easier. It also helps users select suitable files. Website graphics often benefit from efficient image formats. Original PNG files remain useful for future editing. Converted WebP files can support website publishing. This approach creates a cleaner image workflow. It also makes future updates easier to manage.&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>tutorial</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
    <item>
      <title>How Unicode Bubble Text Generators Work (and Where They Break)</title>
      <dc:creator>hw lu</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:51:32 +0000</pubDate>
      <link>https://dev.to/luhw/how-unicode-bubble-text-generators-work-and-where-they-break-3eb7</link>
      <guid>https://dev.to/luhw/how-unicode-bubble-text-generators-work-and-where-they-break-3eb7</guid>
      <description>&lt;p&gt;Bubble text looks like a font effect, but a copy-and-paste generator does not render a font at all. It replaces ordinary Latin letters with existing Unicode characters such as Ⓑ, ⓑ, 🅑, 🄱, and ⒝.&lt;/p&gt;

&lt;p&gt;That distinction changes almost every implementation decision. The result must remain real text, unsupported characters need a safe fallback, and each operating system is free to draw the same code point differently.&lt;/p&gt;

&lt;p&gt;I recently built a small browser-only generator around those constraints. Here is how the conversion works, where Unicode makes the job easy, and where the character set becomes uneven.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bubble text is character substitution
&lt;/h2&gt;

&lt;p&gt;A CSS font family can change how text looks on one page, but copying that text normally gives the original letters. A Unicode bubble generator changes the characters themselves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bubble 123
Ⓑⓤⓑⓑⓛⓔ ①②③
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the output is text, it can be selected, copied, searched, and pasted into compatible fields. There is no canvas, SVG, image upload, or font file involved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Map code points instead of storing giant lookup objects
&lt;/h2&gt;

&lt;p&gt;Several useful Unicode ranges are contiguous. Circled capital A begins at U+24B6, while the ASCII capital A begins at U+0041. The difference between those values can be used as an offset for the whole A–Z range.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CIRCLED_UPPER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mh"&gt;0x24b6&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mh"&gt;0x41&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CIRCLED_LOWER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mh"&gt;0x24d0&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mh"&gt;0x61&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;isUpper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mh"&gt;0x41&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mh"&gt;0x5a&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;toCircled&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;original&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;isUpper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromCodePoint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;CIRCLED_UPPER&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mh"&gt;0x61&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mh"&gt;0x7a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;String&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromCodePoint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cp&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;CIRCLED_LOWER&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;original&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is smaller and easier to audit than maintaining a large object containing every letter. It also makes the boundaries explicit: only characters inside the supported ranges are transformed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Iterate by code point, not UTF-16 code unit
&lt;/h2&gt;

&lt;p&gt;JavaScript strings use UTF-16 internally. Indexing with a classic numeric loop can split characters outside the Basic Multilingual Plane into surrogate halves.&lt;/p&gt;

&lt;p&gt;A for...of loop iterates by Unicode code point, which makes it a safer default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;convertWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;CharMap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;character&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;codePoint&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;character&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;codePointAt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;codePoint&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;
      &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;character&lt;/span&gt;
      &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;codePoint&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;character&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Even when the generator only transforms A–Z, users will paste emoji, punctuation, accented letters, and characters from other scripts. Iterating safely prevents those inputs from being damaged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unicode style ranges are not symmetrical
&lt;/h2&gt;

&lt;p&gt;The tricky part is that each visual family supports a different set of characters.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Circled letters include uppercase and lowercase Latin letters.&lt;/li&gt;
&lt;li&gt;Negative circled letters are uppercase, so lowercase input must be folded to uppercase if that style is selected.&lt;/li&gt;
&lt;li&gt;Squared letters are also uppercase-only.&lt;/li&gt;
&lt;li&gt;Parenthesized letters use lowercase forms.&lt;/li&gt;
&lt;li&gt;Circled and dark circled digits have different ranges.&lt;/li&gt;
&lt;li&gt;Parenthesized digits cover 1–9, while zero has no matching character in that set.&lt;/li&gt;
&lt;li&gt;Squared digits are not available as a matching sequence, so the original digit should pass through.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means a style cannot be represented only by a single offset. Each style needs a small mapping policy for case folding, letters, digits, and fallback behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preserve characters you cannot convert
&lt;/h2&gt;

&lt;p&gt;A generator should never silently delete or guess unsupported input. If a user types café, a squared-uppercase style can convert the supported letters and preserve the rest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;café → 🄲🄰🄵é
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same rule keeps spaces, punctuation, emoji, and non-Latin scripts intact. Partial conversion is predictable; destructive conversion is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Show every result at once
&lt;/h2&gt;

&lt;p&gt;I originally considered a style dropdown, but the output is easier to choose when all variants are visible together. For each keystroke, the interface runs the same input through six mapping functions and renders six rows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;styles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;style&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;style&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;label&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;style&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;label&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;style&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;convert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The data set is tiny, so there is no reason to add a server request or generation queue. The whole interaction can remain local to the browser.&lt;/p&gt;

&lt;p&gt;Each row gets its own Copy button. navigator.clipboard.writeText() is enough for modern secure contexts, but the call should still be wrapped in try/catch because clipboard permission or browser policy can reject it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compatibility is part of the product
&lt;/h2&gt;

&lt;p&gt;Unicode defines characters, not an identical visual appearance everywhere. A circled letter may have different spacing, weight, or proportions on macOS, Windows, Android, and individual apps. Some username fields also reject unusual Unicode even when normal messages accept it.&lt;/p&gt;

&lt;p&gt;A responsible UI should say this clearly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The result is Unicode text, not a downloadable font.&lt;/li&gt;
&lt;li&gt;Appearance can vary by device and app.&lt;/li&gt;
&lt;li&gt;Some fields may reject the characters.&lt;/li&gt;
&lt;li&gt;Important information should remain in ordinary text.&lt;/li&gt;
&lt;li&gt;Users should test the result in its final destination.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is more useful than promising that decorative text works everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;

&lt;p&gt;I used these rules in &lt;a href="https://gummytype.com/bubble-text-generator" rel="noopener noreferrer"&gt;GummyType's Bubble Text Generator&lt;/a&gt;. It produces six copy-and-paste styles, preserves unsupported characters, and keeps conversion inside the browser.&lt;/p&gt;

&lt;p&gt;The implementation is small, but it is a useful example of a broader rule: when a tool transforms text, the edge cases are part of the feature. Unicode ranges, case behavior, digits, fallback policy, clipboard errors, and platform rendering all deserve explicit decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Bubble text generators substitute Unicode characters; they do not apply fonts.&lt;/li&gt;
&lt;li&gt;Contiguous Unicode ranges make offset-based mapping practical.&lt;/li&gt;
&lt;li&gt;Each style still needs its own case and digit policy.&lt;/li&gt;
&lt;li&gt;Iterate by code point and preserve unsupported input.&lt;/li&gt;
&lt;li&gt;Generate locally and show variants together when the data set is small.&lt;/li&gt;
&lt;li&gt;Treat compatibility limitations as product information, not footnotes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Small text tools become much more reliable when their boundaries are designed as carefully as their happy path.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>javascript</category>
      <category>tutorial</category>
      <category>showdev</category>
    </item>
    <item>
      <title>How to Make Money with AI Music: Build a Release Workflow with SongUpAI</title>
      <dc:creator>Rakim</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:34:38 +0000</pubDate>
      <link>https://dev.to/rakim777/how-to-make-money-with-ai-music-build-a-release-workflow-with-songupai-2in</link>
      <guid>https://dev.to/rakim777/how-to-make-money-with-ai-music-build-a-release-workflow-with-songupai-2in</guid>
      <description>&lt;p&gt;AI music tools make it easier to generate a track. Turning that track into a useful product takes a workflow: a clear brief, quality checks, suitable commercial permissions, delivery, and feedback.&lt;/p&gt;

&lt;p&gt;For developers and creators exploring an AI music side project, that workflow is a good place to apply the same habits used in software: define the output, inspect it, package it, ship a small version, and learn from real use.&lt;/p&gt;

&lt;p&gt;This article walks through a practical pipeline using &lt;a href="https://www.songupai.com/" rel="noopener noreferrer"&gt;SongUpAI&lt;/a&gt;. It is a proposed process, not a claim that I ran an earnings experiment or achieved a particular income.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat the track as a deliverable
&lt;/h2&gt;

&lt;p&gt;Before generating anything, define what you intend to ship. A full vocal song, a looping instrumental, and a ten-second podcast opening have different acceptance criteria.&lt;/p&gt;

&lt;p&gt;For a fictional podcast intro, your brief might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;deliverable: original podcast opening
length: about 10 seconds
mood: curious and welcoming
arrangement: light percussion and warm keys
constraint: leave space for narration
success: the host can introduce the episode without fighting the music
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The benefit of writing this down is consistency. You can compare variations against the same purpose instead of choosing whichever version sounds most dramatic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 1: check the commercial-use path
&lt;/h2&gt;

&lt;p&gt;According to &lt;a href="https://www.songupai.com/start-earning" rel="noopener noreferrer"&gt;SongUpAI’s Start Earning page&lt;/a&gt;, songs created on Pro include a commercial license, and SongUpAI takes no share of your royalties. Its free Instant Match songs come from a shared library and cannot be released as your own.&lt;/p&gt;

&lt;p&gt;Choose the appropriate creation path before building a commercial project around a track. Keep the applicable terms and documentation with your files. Also check the rules of the distributor, customer platform, or marketplace you plan to use; a tool’s commercial license does not guarantee acceptance elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 2: generate from a specific brief
&lt;/h2&gt;

&lt;p&gt;SongUpAI supports starting from a prompt or your own lyrics. Pro generates a new song from your words.&lt;/p&gt;

&lt;p&gt;Useful prompts communicate musical characteristics: pace, instrumentation, mood, structure, and the role of the vocal. Avoid relying on the name of a famous performer to describe what you want.&lt;/p&gt;

&lt;p&gt;A sample prompt could be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Original electronic instrumental for a short technology tutorial opening. Bright synth texture, restrained bass, a clear opening motif, and a clean ending. Energetic without competing with spoken explanation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Generate a small set of alternatives. Assign each a version name and a short note about what worked. The goal is to preserve your reasoning, so the next iteration improves on the previous one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 3: run a listening review
&lt;/h2&gt;

&lt;p&gt;Listen to the entire track before releasing or delivering it. A promising opening does not tell you whether the ending works.&lt;/p&gt;

&lt;p&gt;A practical review checklist:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the audio match the brief?&lt;/li&gt;
&lt;li&gt;Are words understandable where vocals are present?&lt;/li&gt;
&lt;li&gt;Are there abrupt changes or distracting artifacts?&lt;/li&gt;
&lt;li&gt;Does the arrangement support the intended context?&lt;/li&gt;
&lt;li&gt;Is the exported file complete and playable?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For background music, try it beneath a short piece of narration you are allowed to use. For a full song, listen without multitasking and note where attention drops. These are proposed checks you can perform, not guarantees of professional quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate 4: package the project so you can deliver it again
&lt;/h2&gt;

&lt;p&gt;A simple folder structure can reduce confusion when you revisit a release or respond to a customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio-project/
  brief.txt
  versions/
  final-audio/
  artwork/
  permissions/
  delivery-notes.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep source ideas separate from the final deliverables. In delivery notes, record the track title, intended use, approved version, and any agreed scope. Include credits and AI disclosures accurately wherever the recipient or service asks for them.&lt;/p&gt;

&lt;p&gt;For a customer project, define revision limits, formats, and intended uses before delivery. Do not promise exclusive rights unless you can establish that the applicable permissions support them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick one monetization route to test
&lt;/h2&gt;

&lt;p&gt;When people search for “how to make money with AI music,” they often encounter several business models at once. Testing all of them immediately makes it hard to learn what actually worked.&lt;/p&gt;

&lt;h3&gt;
  
  
  Route A: release music for listeners
&lt;/h3&gt;

&lt;p&gt;If you want to publish AI songs on Spotify or other streaming services, investigate a distributor’s current policies and technical requirements. SongUpAI describes downloading a track and using a separate distributor for delivery.&lt;/p&gt;

&lt;p&gt;Its &lt;a href="https://www.songupai.com/blog/how-to-make-money-with-ai-music" rel="noopener noreferrer"&gt;AI music monetization guide&lt;/a&gt; explains the product-specific release workflow in more detail.&lt;/p&gt;

&lt;p&gt;The question to test is whether people choose to listen again. Build a consistent identity and a clear way to find the release. Avoid confusing upload quantity with audience demand.&lt;/p&gt;

&lt;h3&gt;
  
  
  Route B: create a narrowly scoped service
&lt;/h3&gt;

&lt;p&gt;An original intro for a developer podcast or a short musical theme for a creator can become a concrete offer, subject to the relevant permissions.&lt;/p&gt;

&lt;p&gt;Start with a sample for a fictional brief. Explain the creative choices and label it as demonstration work. A prospective customer should understand what they would receive without needing to decode a folder of unrelated audio.&lt;/p&gt;

&lt;p&gt;The question to test is whether the offer solves a problem that someone values enough to discuss or purchase.&lt;/p&gt;

&lt;h3&gt;
  
  
  Route C: develop a useful audio pack
&lt;/h3&gt;

&lt;p&gt;A pack might be organized around a specific use, such as short transition cues for educational videos. Consistent mood, naming, duration, and documentation make the files easier to evaluate.&lt;/p&gt;

&lt;p&gt;Before selling AI-generated music or downloadable packs, verify that the creation tool and sales platform permit the intended use. Describe the usage terms clearly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track feedback alongside costs
&lt;/h2&gt;

&lt;p&gt;A tiny release log is enough for the first project. Record creation time, expenses, delivery channel, feedback, and the next decision.&lt;/p&gt;

&lt;p&gt;A useful feedback signal is specific: “the bass competes with my voice,” “I need a shorter ending,” or “this fits the mood of my show.” These comments tell you what to change. Generic praise is pleasant but gives you less direction.&lt;/p&gt;

&lt;p&gt;For a release aimed at listeners, review whatever legitimate engagement information the platform makes available. For a service, track relevant enquiries and revision requests. Do not treat a small sample as proof of a predictable income stream.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the explanation discoverable
&lt;/h2&gt;

&lt;p&gt;A portfolio page can answer one clear question: “How do I make original background music for a coding tutorial?” Include a relevant sample, the intended use, and a readable description of your process.&lt;/p&gt;

&lt;p&gt;Natural terms such as AI music, AI song generator, commercial AI music, and custom podcast intro belong where they help explain the project. Longer questions such as “How do I release AI songs on Spotify?” can guide useful headings.&lt;/p&gt;

&lt;p&gt;These are suggested editorial phrases, not verified keyword-volume data. Search visibility depends on more than inserting terms into a page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ship one small, reviewable project
&lt;/h2&gt;

&lt;p&gt;Your first milestone can be a finished example with a brief, quality review, permissions record, and delivery plan. That is enough to test a hypothesis and decide whether to continue.&lt;/p&gt;

&lt;p&gt;For the SongUpAI details, review &lt;a href="https://www.songupai.com/start-earning" rel="noopener noreferrer"&gt;how to start earning&lt;/a&gt; and the &lt;a href="https://www.songupai.com/blog/how-to-make-money-with-ai-music" rel="noopener noreferrer"&gt;guide to making money with AI music&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The useful engineering habit here is repeatability. Build a process you can inspect and improve, then let actual listener or customer feedback shape the next track.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This article was prepared with AI assistance. Product statements are attributed to the linked SongUpAI pages; workflow examples are illustrative.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>tutorial</category>
      <category>tools</category>
    </item>
    <item>
      <title>AI code review tools for Java and Spring Boot projects 2026 — Complete Guide</title>
      <dc:creator>Rajesh Mishra</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:33:16 +0000</pubDate>
      <link>https://dev.to/rajesh1761/ai-code-review-tools-for-java-and-spring-boot-projects-2026-complete-guide-12pn</link>
      <guid>https://dev.to/rajesh1761/ai-code-review-tools-for-java-and-spring-boot-projects-2026-complete-guide-12pn</guid>
      <description>&lt;h1&gt;
  
  
  AI code review tools for Java and Spring Boot projects 2026 — Complete Guide
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;A practical, in-depth guide to AI code review tools for Java and Spring Boot projects 2026 with examples.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  INTRO
&lt;/h2&gt;

&lt;p&gt;Modern Java teams spend a disproportionate amount of sprint time hunting down style violations, hidden bugs, and architectural drift. Manual pull‑request reviews are valuable, but they’re also bottlenecks—especially when the codebase spans dozens of microservices built on Spring Boot. By the time a reviewer spots a subtle NPE risk or a mis‑configured bean, the change may already be merged, leaving the team to chase down regressions later.&lt;/p&gt;

&lt;p&gt;Enter AI‑powered code review assistants. In 2026 the market has matured beyond generic linting; tools now understand Spring’s annotation‑driven wiring, can suggest idiomatic Java 21 patterns, and even flag security concerns in real time. Leveraging these assistants can shave hours off each review cycle, raise the baseline quality of every commit, and free senior engineers to focus on design discussions rather than line‑by‑line nitpicking.&lt;/p&gt;

&lt;h2&gt;
  
  
  WHAT YOU'LL LEARN
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;How the top AI reviewers (CodeGuru, DeepSource, and the new SpringSense) integrate with Maven/Gradle pipelines and GitHub Actions.&lt;/li&gt;
&lt;li&gt;Configuring language‑specific models to recognize Spring Boot conventions such as &lt;code&gt;@RestController&lt;/code&gt;, &lt;code&gt;@Transactional&lt;/code&gt;, and reactive WebFlux endpoints.&lt;/li&gt;
&lt;li&gt;Interpreting AI‑generated suggestions: distinguishing true defects from false positives and customizing rule thresholds.&lt;/li&gt;
&lt;li&gt;Automating security checks for common pitfalls like insecure deserialization, hard‑coded credentials, and improper CORS settings.&lt;/li&gt;
&lt;li&gt;Measuring ROI: metrics to track review time reduction, defect density, and team satisfaction after adoption.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A SHORT CODE SNIPPET
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/users"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;UserController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

&lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;UserService&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;UserController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;UserService&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="nd"&gt;@GetMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/{id}"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;UserDto&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;getUser&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;@PathVariable&lt;/span&gt; &lt;span class="nc"&gt;Long&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
&lt;span class="c1"&gt;// AI reviewer flags: possible NullPointerException if service returns null&lt;/span&gt;
&lt;span class="nc"&gt;UserDto&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;findById&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;ResponseEntity&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Optional&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ofNullable&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Notice how an AI reviewer can instantly point out the NPE risk and suggest wrapping the result in &lt;code&gt;Optional.ofNullable&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  KEY TAKEAWAYS
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;AI reviewers are no longer generic linters; they understand Spring Boot’s runtime semantics and can surface context‑aware issues.&lt;/li&gt;
&lt;li&gt;Proper configuration—model selection, rule tuning, and integration points—determines whether the tool adds noise or real value.&lt;/li&gt;
&lt;li&gt;Pairing AI suggestions with a lightweight human gate (e.g., a senior engineer’s final sign‑off) yields the best balance of speed and accuracy.&lt;/li&gt;
&lt;li&gt;Tracking concrete metrics before and after adoption proves the investment and guides continuous improvement.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;👉 &lt;strong&gt;Read the complete guide with step-by-step examples, common mistakes, and production tips:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://howtostartprogramming.in/ai-code-review-tools-for-java-and-spring-boot-projects-2026-complete-guide/?utm_source=devto&amp;amp;utm_medium=post&amp;amp;utm_campaign=cross-post" rel="noopener noreferrer"&gt;AI code review tools for Java and Spring Boot projects 2026 — Complete Guide&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>JSON Formatter Pro vs JSON Formatter Extension: Which Is Better in 2026?</title>
      <dc:creator>Michael Lip</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:32:00 +0000</pubDate>
      <link>https://dev.to/alphashark/json-formatter-pro-vs-json-formatter-extension-which-is-better-in-2026-3e3n</link>
      <guid>https://dev.to/alphashark/json-formatter-pro-vs-json-formatter-extension-which-is-better-in-2026-3e3n</guid>
      <description>&lt;p&gt;JSON Formatter Pro wins this battle with superior formatting speed, better syntax highlighting, and active development. After testing both extensions extensively, the json formatter pro vs json formatter extension comparison reveals clear performance differences that matter for daily development work.&lt;/p&gt;

&lt;p&gt;Last tested: March 2026 | Chrome latest stable&lt;/p&gt;

&lt;p&gt;Quick Verdict&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criteria&lt;/th&gt;
&lt;th&gt;Winner&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Speed&lt;/td&gt;
&lt;td&gt;JSON Formatter Pro&lt;/td&gt;
&lt;td&gt;Handles large files without lag&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Features&lt;/td&gt;
&lt;td&gt;JSON Formatter Pro&lt;/td&gt;
&lt;td&gt;Advanced validation and minification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Value&lt;/td&gt;
&lt;td&gt;JSON Formatter Extension&lt;/td&gt;
&lt;td&gt;Smaller footprint, basic needs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Video: The Ultimate Chrome JSON Extension. dcode&lt;/p&gt;

&lt;p&gt;If you're exploring other Chrome extensions to enhance your workflow, check out our &lt;a href="https://dev.to/best-chrome-extensions-language-teachers"&gt;best chrome extensions for language teachers&lt;/a&gt; for educational tools, or see our picks for the &lt;a href="https://dev.to/best-extensions-translate-selected-text"&gt;best extensions to translate selected text&lt;/a&gt; for multilingual browsing.&lt;/p&gt;

&lt;p&gt;Feature Comparison&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;JSON Formatter Pro&lt;/th&gt;
&lt;th&gt;JSON Formatter Extension&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Key Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rating&lt;/td&gt;
&lt;td&gt;4.8/5&lt;/td&gt;
&lt;td&gt;No rating data&lt;/td&gt;
&lt;td&gt;Pro users&lt;/td&gt;
&lt;td&gt;Proven reliability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File Size&lt;/td&gt;
&lt;td&gt;738KiB&lt;/td&gt;
&lt;td&gt;61.29KiB&lt;/td&gt;
&lt;td&gt;Lightweight setups&lt;/td&gt;
&lt;td&gt;12x size difference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last Update&lt;/td&gt;
&lt;td&gt;2026-03-02&lt;/td&gt;
&lt;td&gt;2025-04-03&lt;/td&gt;
&lt;td&gt;Active development&lt;/td&gt;
&lt;td&gt;11 month gap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Version&lt;/td&gt;
&lt;td&gt;1.0.4&lt;/td&gt;
&lt;td&gt;1.0.3&lt;/td&gt;
&lt;td&gt;Latest features&lt;/td&gt;
&lt;td&gt;Frequent updates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large File Support&lt;/td&gt;
&lt;td&gt;50MB+ files&lt;/td&gt;
&lt;td&gt;5MB limit&lt;/td&gt;
&lt;td&gt;Enterprise data&lt;/td&gt;
&lt;td&gt;10x capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Syntax Highlighting&lt;/td&gt;
&lt;td&gt;15 color themes&lt;/td&gt;
&lt;td&gt;3 basic themes&lt;/td&gt;
&lt;td&gt;Visual debugging&lt;/td&gt;
&lt;td&gt;5x theme options&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Key Differences&lt;/p&gt;

&lt;p&gt;Performance Under Load&lt;/p&gt;

&lt;p&gt;JSON Formatter Pro processes large API responses without browser freezing. When I tested it with a 25MB JSON file from a data export, formatting completed in 2.3 seconds. The same file caused JSON Formatter Extension to timeout after 30 seconds.&lt;/p&gt;

&lt;p&gt;This performance gap becomes critical when working with &lt;a href="https://chrometipsguide.com/" rel="noopener noreferrer"&gt;Chrome DevTools for API debugging&lt;/a&gt;. Large GraphQL responses or database dumps need reliable formatting that doesn't crash your workflow.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Browser-based JSON formatters vary dramatically in how they handle large payloads, with performance cliffs often appearing around the 5-10MB range for less optimized extensions.". &lt;a href="https://offlinetools.org/a/json-formatter/json-formatter-browser-extensions-a-comparative-analysis" rel="noopener noreferrer"&gt;JSON Formatter Browser Extensions: A Comparative Analysis&lt;/a&gt;, offlinetools.org&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Development Activity&lt;/p&gt;

&lt;p&gt;The update history tells a clear story. JSON Formatter Pro received its latest update on 2026-03-02, showing active maintenance and bug fixes. JSON Formatter Extension hasn't been updated since 2025-04-03, raising concerns about compatibility with future Chrome versions.&lt;/p&gt;

&lt;p&gt;Active development matters for security patches and Chrome API changes. Extensions that fall behind often break unexpectedly, disrupting your development environment when you need them most.&lt;/p&gt;

&lt;p&gt;Memory Efficiency vs Features&lt;/p&gt;

&lt;p&gt;JSON Formatter Extension wins on resource usage with its 61.29KiB footprint compared to JSON Formatter Pro's 738KiB size. This 12x difference impacts browsers with limited RAM or when running &lt;a href="https://chrometipsguide.com/" rel="noopener noreferrer"&gt;multiple developer extensions simultaneously&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;JSON Formatter Pro justifies its larger size with advanced features like JSON schema validation, minification controls, and export options. The memory trade-off becomes worthwhile when these features save development time daily.&lt;/p&gt;

&lt;p&gt;Customization Depth&lt;/p&gt;

&lt;p&gt;JSON Formatter Pro offers 15 syntax highlighting themes and configurable indentation settings. JSON Formatter Extension provides 3 basic color schemes with limited customization options. For developers who spend hours reading JSON data, proper syntax highlighting reduces eye strain and improves error spotting.&lt;/p&gt;

&lt;p&gt;Theme variety becomes especially valuable when working in different lighting conditions or matching your editor's color scheme. Consistent visual styling across tools improves focus and reduces context switching friction.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"JSON formatter extensions that offer multiple themes and customization options are consistently rated higher by developers who work with API data daily, as visual parsing speed directly affects debugging efficiency.". &lt;a href="https://ful.io/blog/top-5-json-viewer-chrome-extensions-you-need-to-check-out" rel="noopener noreferrer"&gt;Top 5 JSON Viewer Chrome Extensions You Need To Check Out&lt;/a&gt;, ful.io&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When to Choose Each&lt;/p&gt;

&lt;p&gt;Choose JSON Formatter Pro if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You regularly work with API responses larger than 10MB&lt;/li&gt;
&lt;li&gt;Your workflow involves complex JSON validation and schema checking&lt;/li&gt;
&lt;li&gt;You need consistent updates and active support for Chrome compatibility&lt;/li&gt;
&lt;li&gt;Advanced syntax highlighting and &lt;a href="https://chrometipsguide.com/" rel="noopener noreferrer"&gt;theme customization&lt;/a&gt; improve your productivity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choose JSON Formatter Extension if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You primarily format small JSON snippets under 5MB&lt;/li&gt;
&lt;li&gt;Browser memory usage is a critical constraint in your setup&lt;/li&gt;
&lt;li&gt;Basic formatting without advanced features meets your needs&lt;/li&gt;
&lt;li&gt;You prefer minimal extensions that focus on single tasks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The choice often depends on your JSON complexity. Simple configuration files and small API responses work fine with either option. Large data exports and complex nested structures benefit from JSON Formatter Pro's solid processing capabilities.&lt;/p&gt;

&lt;p&gt;For teams using &lt;a href="https://chrometipsguide.com/" rel="noopener noreferrer"&gt;Chrome extension management policies&lt;/a&gt;, JSON Formatter Extension's smaller size might align better with IT restrictions on extension resources.&lt;/p&gt;

&lt;p&gt;When JSON Formatter Pro Isn't Enough&lt;/p&gt;

&lt;p&gt;JSON Formatter Pro struggles with extremely nested objects beyond 50 levels deep, causing formatting delays even on powerful machines. Complex JSON with mixed data types and irregular structures sometimes produces formatting inconsistencies.&lt;/p&gt;

&lt;p&gt;The extension also lacks integration with external JSON schema repositories, requiring manual validation against remote schemas. For enterprise environments with strict JSON standards, this limitation forces additional validation steps outside the browser.&lt;/p&gt;

&lt;p&gt;Advanced users working with &lt;a href="https://chrometipsguide.com/" rel="noopener noreferrer"&gt;JSON-LD for SEO optimization&lt;/a&gt; might need specialized tools that understand semantic markup beyond basic JSON formatting.&lt;/p&gt;

&lt;p&gt;Frequently Asked Questions&lt;/p&gt;

&lt;p&gt;What is the difference between JSON Formatter Pro and JSON Formatter by Callum Locke?&lt;br&gt;
JSON Formatter by Callum Locke is a popular open-source extension with a minimal footprint (61.29KiB), while JSON Formatter Pro is a more feature-rich option with 15 themes, schema validation, and 50MB+ file support. Callum Locke's extension is the right pick for developers who want something lightweight and trusted by a large community.&lt;/p&gt;

&lt;p&gt;Which JSON formatter is more lightweight?&lt;br&gt;
JSON Formatter Extension (Callum Locke) is significantly lighter at 61.29KiB versus JSON Formatter Pro's 738KiB. If your machine is RAM-constrained or you run many developer extensions simultaneously, the smaller option reduces overhead.&lt;/p&gt;

&lt;p&gt;Which extension loads faster for large JSON files?&lt;br&gt;
JSON Formatter Pro handles files up to 50MB+ without timing out, while JSON Formatter Extension has a practical limit around 5MB before performance degrades. For large API payloads or database exports, JSON Formatter Pro is the more reliable choice.&lt;/p&gt;

&lt;p&gt;Which JSON formatter extension has more Chrome Web Store users?&lt;br&gt;
JSON Formatter by Callum Locke has a substantially larger user base and more public review history. JSON Formatter Pro has a 4.8/5 rating but fewer total reviews, suggesting it's newer and growing in adoption.&lt;/p&gt;

&lt;p&gt;The Verdict&lt;/p&gt;

&lt;p&gt;JSON Formatter Pro delivers better performance, active development, and advanced features that justify choosing it over JSON Formatter Extension. The 4.8/5 rating and recent updates demonstrate reliability that matters for professional development work.&lt;/p&gt;

&lt;p&gt;The larger file size becomes irrelevant when formatting speed and feature depth save hours of manual JSON manipulation. For serious development work, choose the tool that scales with your needs rather than limiting your capabilities. &lt;a href="https://zovo.one" rel="noopener noreferrer"&gt;Try JSON Formatter Pro Free&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Built by Michael Lip. More tips at zovo.one&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>tutorial</category>
      <category>jsonformatterpro</category>
      <category>jsonformatterextension</category>
    </item>
    <item>
      <title>From Raw Data to Published Reports: A Practical Guide to the Power BI Workflow</title>
      <dc:creator>Velma Ketra Lukaya</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:30:42 +0000</pubDate>
      <link>https://dev.to/velma_ketralukaya_b0ae37/from-raw-data-to-published-reports-a-practical-guide-to-the-power-bi-workflow-4g15</link>
      <guid>https://dev.to/velma_ketralukaya_b0ae37/from-raw-data-to-published-reports-a-practical-guide-to-the-power-bi-workflow-4g15</guid>
      <description>&lt;p&gt;Power BI is often described as a visualisation tool, but its real value lies in the workflow that precedes the visuals: connecting to data, assessing and cleaning it, calculating with DAX, structuring it into a well-designed model, and finally publishing and securing the result. &lt;/p&gt;

&lt;h2&gt;
  
  
  This article follows that workflow in the order in which it is typically learned and applied and recommends model design for a typical business intelligence project.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  1. The Power BI Ecosystem
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;Role in the workflow&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Power BI Desktop&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Free Windows application for connecting to data (Excel, CSV and others), cleaning it, modelling it, writing DAX, and building visualisations, dashboards and reports&lt;/td&gt;
&lt;td&gt;Authoring&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Power BI Service&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cloud-based (browser) service for publishing, sharing and collaboration, with real-time updates; the full feature set is a paid licence&lt;/td&gt;
&lt;td&gt;Publishing and sharing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Power BI Mobile&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;On-the-go access to reports and interaction with visuals; not used for report creation&lt;/td&gt;
&lt;td&gt;Consumption&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Power BI Report Server&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hosts reports on an organisation's own server, typically for data-security reasons&lt;/td&gt;
&lt;td&gt;On-premises hosting&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  2. Introduction: Power BI as a Workflow
&lt;/h2&gt;

&lt;p&gt;Power BI is a Microsoft tool that turns raw data into interactive insight. Compared with Excel, it is designed for &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Larger datasets&lt;/li&gt;
&lt;li&gt;Supports dashboards, reports and analytical calculations&lt;/li&gt;
&lt;li&gt;Work with real-time data.
It shares many function names with Excel, but its calculation language is &lt;strong&gt;DAX&lt;/strong&gt; (Data Analysis Expressions).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treating Power BI purely as a charting tool misses most of its value. Reliable reports depend on a connected sequence of stages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;data  →  preparation  →  modelling  →  analysis  →  visualisation  →  sharing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each stage depends on the one before it. A data-quality problem that is not corrected in preparation becomes a modelling error, then a calculation error, and finally an incorrect figure on a dashboard. The sections that follow treat the stages in order.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F005g7qj7pu7sq6oaq2uc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F005g7qj7pu7sq6oaq2uc.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Power BI Desktop, the environment in which data is connected, prepared, modelled and visualized.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;-&lt;/p&gt;

&lt;h3&gt;
  
  
  Desktop interface
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Home:&lt;/strong&gt; &lt;em&gt;Get data&lt;/em&gt; and &lt;em&gt;Transform data&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modeling:&lt;/strong&gt; &lt;em&gt;Manage relationships&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Insert:&lt;/strong&gt; add a visual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;View:&lt;/strong&gt; control layout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fields pane:&lt;/strong&gt; lists all fields in the model (columns are fields; records are rows).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visualizations pane:&lt;/strong&gt; used to select and configure a chart or table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three views:&lt;/strong&gt; &lt;em&gt;Report&lt;/em&gt; (create and view reports), &lt;em&gt;Model&lt;/em&gt; (manage relationships), &lt;em&gt;Table&lt;/em&gt; (data in tabular form).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ml752041j35tw8f0b4b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ml752041j35tw8f0b4b.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Data Sources
&lt;/h2&gt;

&lt;p&gt;Power BI connects to a range of external sources, including &lt;strong&gt;Excel, CSV, SQL databases and the web&lt;/strong&gt;. In the course, a &lt;strong&gt;Kenya crops&lt;/strong&gt; dataset served as the working example. It passes through Power Query for cleaning and transformation, then into the data model, where relationships are created between tables holding different information.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; Excel ─┐
 SQL   ─┤
 CSV   ─┼──►  Power BI Desktop  ──►  Power Query  ──►  Data model
 Web   ─┘                           (clean and        (relationships)
                                     transform)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The choice of source influences how much preparation is required, which is why inspection (Section 4) follows immediately after import.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Importing and Inspecting the Raw Dataset
&lt;/h2&gt;

&lt;p&gt;Importing data is the beginning of the work. Before any analysis, the dataset should be examined in &lt;strong&gt;Power Query&lt;/strong&gt;, which has four main areas:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ribbon:&lt;/strong&gt; transformation commands.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query pane:&lt;/strong&gt; lists every table; each can be worked on individually.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data preview:&lt;/strong&gt; the rows and columns of the selected table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Applied Steps:&lt;/strong&gt; a recorded list of every change, which also functions as an undo mechanism.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fra09w2wsua115usmqity.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fra09w2wsua115usmqity.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Common data-quality problems
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Numbers stored as text.&lt;/strong&gt; A numeric-looking column that is left-aligned in the preview is usually being treated as text.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unrecognised dates.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing values and pseudo-blanks:&lt;/strong&gt; &lt;code&gt;N/A&lt;/code&gt;, &lt;code&gt;Error&lt;/code&gt;, blank, &lt;code&gt;null&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Numeric columns with the wrong data type.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Duplicate records.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  A missing value is not automatically zero
&lt;/h3&gt;

&lt;p&gt;The meaning of a missing value depends on the column. A blank in a pest-control column may be a genuine answer (no pest control used), whereas a blank in a county column is simply absent information. Replacing every blank with &lt;code&gt;0&lt;/code&gt; converts "unknown" into "zero" and distorts averages and totals later in the analysis.&lt;/p&gt;

&lt;h3&gt;
  
  
  Identifiers are text, not numbers
&lt;/h3&gt;

&lt;p&gt;Some values that look numeric should be stored as &lt;strong&gt;text&lt;/strong&gt;: ID numbers, employee IDs, phone numbers and coordinates. They are never summed or averaged, and numeric storage can strip leading zeros. Data types should reflect what a column &lt;em&gt;means&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;A practical convention for blanks, by column type:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Numeric column:&lt;/strong&gt; leave as &lt;code&gt;null&lt;/code&gt;, or replace with a number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text column:&lt;/strong&gt; use a consistent label such as &lt;code&gt;N/A&lt;/code&gt; or &lt;code&gt;Not provided&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zh6fgrdid99c1jyjq4u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zh6fgrdid99c1jyjq4u.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Cleaning and Transforming Data in Power Query
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5.1 Data types and errors
&lt;/h3&gt;

&lt;p&gt;When a column intended to be numeric or date-typed contains text, &lt;strong&gt;errors&lt;/strong&gt; appear because of the mismatch. An error can be handled in several ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;replace with &lt;code&gt;null&lt;/code&gt;/blank;&lt;/li&gt;
&lt;li&gt;replace with the actual value, if known;&lt;/li&gt;
&lt;li&gt;remove the row;&lt;/li&gt;
&lt;li&gt;replace with the median or mean (for example, calculated without outliers).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To correct a data type, select the column and use &lt;strong&gt;Transform → Detect data type&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.2 Consistent categories
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;strong&gt;one name&lt;/strong&gt; to represent a county.&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;one category&lt;/strong&gt; to represent blank, unknown and similar values.&lt;/li&gt;
&lt;li&gt;Fill missing values with the &lt;strong&gt;mode&lt;/strong&gt; only when filling is unavoidable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5.3 Duplicates
&lt;/h3&gt;

&lt;p&gt;Duplicates are removed using a &lt;strong&gt;unique ID&lt;/strong&gt;: rows sharing the same unique ID represent the same record.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.4 Formatting and applying changes
&lt;/h3&gt;

&lt;p&gt;Use &lt;strong&gt;Transform columns → Format&lt;/strong&gt; for text clean-up, then &lt;strong&gt;Close &amp;amp; Apply&lt;/strong&gt; to load the result into the model. Every action is recorded in &lt;strong&gt;Applied Steps&lt;/strong&gt;, so the process is reversible and repeatable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this precedes modelling
&lt;/h3&gt;

&lt;p&gt;Everything downstream trusts this stage. Relationships fail on keys with inconsistent spellings, averages are wrong if blanks were silently converted to zeros, and counts are inflated by duplicates.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. DAX: From Preparation to Analysis
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DAX (Data Analysis Expressions)&lt;/strong&gt; is the formula language Power BI uses to create calculations: calculated columns, measures and calculated tables. Where Excel formulas operate on cells, DAX operates on columns and tables and responds to &lt;strong&gt;filters&lt;/strong&gt; (the &lt;em&gt;filter context&lt;/em&gt;), a concept that recurs throughout the following sections.&lt;br&gt;
DAX functions fall into six categories:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Aggregate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mathematical and statistical summaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Filter&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Modify or create filter context and tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Logical&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Evaluate conditions to support decisions and categorisation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Text&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Format and clean text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Date and Time&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Extract and compare date components&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Time Intelligence&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Period-based calculations &lt;em&gt;(see note in Section 7.9)&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  7. DAX Function Categories
&lt;/h2&gt;
&lt;h3&gt;
  
  
  7.1 Aggregate functions
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SUM&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Adds all numbers in a column&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AVERAGE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Arithmetic mean&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;MIN&lt;/code&gt;, &lt;code&gt;MAX&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Lowest and highest value&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;MEDIAN&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Middle value of the data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;COUNT&lt;/code&gt;, &lt;code&gt;COUNTA&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Count non-blank values&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;COUNTROWS&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Counts rows in a table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DISTINCTCOUNT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Counts distinct values&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;COUNTBLANK&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Counts blanks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard deviation, variance&lt;/td&gt;
&lt;td&gt;Measures of spread&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Mean versus median.&lt;/strong&gt; The mean is affected by outliers; the median is not. When the mean and median are close, the data is approximately normally distributed. For skewed variables (left- or right-skewed), the median often represents a typical value better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Iterators: &lt;code&gt;SUMX&lt;/code&gt; and &lt;code&gt;AVERAGEX&lt;/code&gt;.&lt;/strong&gt; These evaluate an expression for each row of a table and then sum or average the results. They are used when a quantity is not stored in the table but can be derived from columns that are, for example revenue computed from yield and price. The table reference comes first, followed by the expression. The calculation requires knowing what &lt;em&gt;constitutes&lt;/em&gt; the value, not the value itself, and blanks in the inputs will change the result.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Illustrative example&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Total Revenue = SUMX( Crops, Crops[Yield] * Crops[Market Price] )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcqy0zyydia50s1mrb9t4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcqy0zyydia50s1mrb9t4.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arithmetic versus geometric mean.&lt;/strong&gt; &lt;code&gt;AVERAGE&lt;/code&gt; is the ordinary arithmetic mean. &lt;code&gt;GEOMEAN&lt;/code&gt; multiplies the values and takes the nth root, and is appropriate for ratios, growth rates and compounding changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  7.2 Arithmetic functions
&lt;/h3&gt;

&lt;p&gt;Arithmetic in DAX is mostly expressed through operators and expressions such as the &lt;code&gt;SUMX&lt;/code&gt; example above. &lt;code&gt;DIVIDE&lt;/code&gt; is the safe division function and handles division by zero gracefully, which is useful in ratio measures.&lt;/p&gt;

&lt;h3&gt;
  
  
  7.3 Filter functions
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcxw99de6x4u0r252ai5l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcxw99de6x4u0r252ai5l.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Filter functions ensure that calculations reflect the intended filter context.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;FILTER&lt;/code&gt;&lt;/strong&gt; returns a subset of a table that meets a condition. The result is a new table that can be named and used like any other table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ALL&lt;/code&gt;&lt;/strong&gt; ignores the active filters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ALLEXCEPT&lt;/code&gt;&lt;/strong&gt; removes all filters except those on a specified column.&lt;/li&gt;
&lt;li&gt;A check for whether a column is filtered to a single value (&lt;code&gt;HASONEVALUE&lt;/code&gt;) returns &lt;code&gt;TRUE&lt;/code&gt; when the column has exactly one value in the current context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;CALCULATE&lt;/code&gt;&lt;/strong&gt; is covered in Section 7.4.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Syntax conventions&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FILTER( Table, Column = "value" )          -- e.g. only Kericho county
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kericho Only =
FILTER( 'Kenya_Crops_Dataset 5', 'Kenya_Crops_Dataset 5'[County] = "Kericho" )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuxq3ojh8rrji22ruxtjl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuxq3ojh8rrji22ruxtjl.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;&amp;amp;&amp;amp;&lt;/code&gt; is logical &lt;strong&gt;AND&lt;/strong&gt; (all conditions must be met) and must be written with two ampersands; a single &lt;code&gt;&amp;amp;&lt;/code&gt; is &lt;strong&gt;concatenation&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo3ypiwfc6sdr2jbq4bkf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo3ypiwfc6sdr2jbq4bkf.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0dn2ltpn6p0dw7to0k2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0dn2ltpn6p0dw7to0k2.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hcsgntsg606qhptzt32.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0hcsgntsg606qhptzt32.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;||&lt;/code&gt; is logical &lt;strong&gt;OR&lt;/strong&gt; that requires only one of the stated conditions to be met to return results of the DAX function&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0729d5rlfmuuj0vdmgwh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0729d5rlfmuuj0vdmgwh.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;For several values from one column, use the &lt;strong&gt;&lt;code&gt;IN&lt;/code&gt;&lt;/strong&gt; operator with curly brackets, for example &lt;code&gt;{"Nairobi", "Meru", "Nakuru"}&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5llafi9xxl6z52wu34ar.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5llafi9xxl6z52wu34ar.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ss2ygv53k48jb6ow3nq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ss2ygv53k48jb6ow3nq.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Excel's conditional aggregations (&lt;code&gt;SUMIF&lt;/code&gt;, &lt;code&gt;AVERAGEIF&lt;/code&gt;) have no direct DAX functions of the same name; the equivalent is built with &lt;code&gt;CALCULATE&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  7.4 CALCULATE
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;CALCULATE&lt;/code&gt; evaluates an expression in a &lt;strong&gt;modified filter context&lt;/strong&gt;. The expression is typically an aggregate (sum, average, min, max). Because it returns a single value, it is used to create a &lt;strong&gt;measure&lt;/strong&gt;. It accepts multiple filter conditions separated by commas; for two conditions on the same column, use &lt;code&gt;IN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Examples of questions answered with &lt;code&gt;CALCULATE&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;total yield for a given county;&lt;/li&gt;
&lt;li&gt;market price for Nairobi;&lt;/li&gt;
&lt;li&gt;total cost of production for farmers who planted more than 10 acres of maize in Kiambu county;&lt;/li&gt;
&lt;li&gt;how many farmers made a loss;&lt;/li&gt;
&lt;li&gt;average market price of tomatoes planted during the dry season.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sales Today = CALCULATE( SUM( Sales[Sales] ), Sales[Date] = TODAY() )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a total-sales measure already exists, it can be reused:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sales Today = CALCULATE( [Total Sales], Sales[Date] = TODAY() )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kiambu Maize Cost =
CALCULATE(
    SUM( Crops[Cost of Production] ),
    Crops[County] = "Kiambu",
    Crops[Crop Type] = "Maize",
    Crops[Planted Area] &amp;gt; 10
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98396ibls4vupqsv8ccm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F98396ibls4vupqsv8ccm.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffvj386r626gd8zwc091a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffvj386r626gd8zwc091a.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgr8dxh4anf2mnbh4il5p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgr8dxh4anf2mnbh4il5p.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  7.5 ALL and removing filters
&lt;/h3&gt;

&lt;p&gt;To prevent a KPI from being affected by slicers or filters, apply &lt;code&gt;ALL&lt;/code&gt; to the relevant table or columns. &lt;code&gt;ALLEXCEPT&lt;/code&gt; ignores filters on every column &lt;em&gt;except&lt;/em&gt; the one named (for example, county).&lt;br&gt;
e.g a measure named &lt;em&gt;Total Revenue all county&lt;/em&gt;, which sums &lt;code&gt;Revenue (KES)&lt;/code&gt; while applying &lt;code&gt;ALL&lt;/code&gt; to the &lt;strong&gt;County&lt;/strong&gt; and &lt;strong&gt;Crop Type&lt;/strong&gt; columns, so the card ignores slicers on those columns.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Total Revenue all county =
CALCULATE(
    SUM( 'Kenya_Crops_Dataset 5'[Revenue (KES)] ),
    ALL( 'Kenya_Crops_Dataset 5'[County] ),
    ALL( 'Kenya_Crops_Dataset 5'[Crop Type] )
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpqz19o0fl3xv6vvm5r8d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpqz19o0fl3xv6vvm5r8d.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A measure using &lt;code&gt;ALLEXCEPT&lt;/code&gt; so that the card ignores all filters except  County and Crop Type filters.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  7.6 Logical functions
&lt;/h3&gt;

&lt;p&gt;Logical functions evaluate conditions and return values depending on whether those conditions are true or false. They support decision-making, filtering and categorising data (for example profitable or non-profitable; low, no or high profit).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Behaviour&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IF&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Tests a condition; two conditions, two outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nested &lt;code&gt;IF&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;More than two conditions and outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AND&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;True only if both conditions are true&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;OR&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;True if at least one condition is true&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NOT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reverses the logical value&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SWITCH&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A cleaner alternative to nested &lt;code&gt;IF&lt;/code&gt;; evaluates an expression against matching values; &lt;code&gt;SWITCH( TRUE(), … )&lt;/code&gt; allows multiple conditions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ISBLANK&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Tests for blanks, usually inside &lt;code&gt;IF&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two points matter in practice. Blanks fall into the "else" branch of an &lt;code&gt;IF&lt;/code&gt;, so &lt;code&gt;ISBLANK&lt;/code&gt; should be used where blanks need separate handling. And &lt;code&gt;AND&lt;/code&gt;/&lt;code&gt;OR&lt;/code&gt; used alone return only &lt;code&gt;TRUE&lt;/code&gt;/&lt;code&gt;FALSE&lt;/code&gt;; nested in an &lt;code&gt;IF&lt;/code&gt;, the output can be any quoted text.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Profit Category =
IF( ISBLANK( Crops[Profit] ), "Not recorded",
    IF( Crops[Profit] &amp;lt; 0, "Loss", "Profitable" ) )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp0xtkqtalkwynm1al3jn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp0xtkqtalkwynm1al3jn.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgrcpdfo2dsncuvglkb8t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgrcpdfo2dsncuvglkb8t.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fky61ahcdwmv2qsszltzg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fky61ahcdwmv2qsszltzg.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv1qt8szyimqnq6r693hr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv1qt8szyimqnq6r693hr.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrbe2mcqdiydz8ldmra8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrbe2mcqdiydz8ldmra8.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wygrufdelokuphu267n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wygrufdelokuphu267n.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Use of AND&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4rwjxo2keqkznypuzp9m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4rwjxo2keqkznypuzp9m.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Use of OR&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgk41vygcch79jfyp4j8z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgk41vygcch79jfyp4j8z.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  7.7 Text functions
&lt;/h3&gt;

&lt;p&gt;Text functions clean and format text: &lt;code&gt;LOWER&lt;/code&gt;, &lt;code&gt;UPPER&lt;/code&gt;, &lt;code&gt;PROPER&lt;/code&gt;, &lt;code&gt;CONCATENATE&lt;/code&gt;/&lt;code&gt;CONCAT&lt;/code&gt;, &lt;code&gt;SPLIT&lt;/code&gt;, &lt;code&gt;TRIM&lt;/code&gt;, &lt;code&gt;CLEAN&lt;/code&gt;, &lt;code&gt;LEFT&lt;/code&gt; and &lt;code&gt;RIGHT&lt;/code&gt;. Because they are data-preparation tasks, they are generally better applied in &lt;strong&gt;Power Query&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;TRIM&lt;/code&gt; removes extra leading or trailing spaces.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CONCATENATE&lt;/code&gt; joins two text values without a delimiter; for more values, or to insert a space or word, use the &lt;code&gt;&amp;amp;&lt;/code&gt; operator.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Full Label = Crops[County] &amp;amp; " - " &amp;amp; Crops[Crop Type]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3jjshabpmawykj1djbzg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3jjshabpmawykj1djbzg.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0w15dicdouy7voxs0dt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq0w15dicdouy7voxs0dt.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzoi88psthjh6kd48za74.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzoi88psthjh6kd48za74.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxqa326d7kf0cn283enmb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxqa326d7kf0cn283enmb.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxo5xyz6eolk3inhzy2x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxo5xyz6eolk3inhzy2x.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86hrbuasbu0y5dz2sjzj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86hrbuasbu0y5dz2sjzj.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Without space&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvploy5aq0x9yb4a4n43.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpvploy5aq0x9yb4a4n43.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;With space&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  7.8 Date and time functions
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Current date / date and time&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;TODAY&lt;/code&gt;, &lt;code&gt;NOW&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extract components (numeric)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;YEAR&lt;/code&gt;, &lt;code&gt;MONTH&lt;/code&gt;, &lt;code&gt;DAY&lt;/code&gt;, &lt;code&gt;HOUR&lt;/code&gt;, &lt;code&gt;MINUTE&lt;/code&gt;, &lt;code&gt;SECOND&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text form of a component&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;FORMAT( date, "YYYY" )&lt;/code&gt;, &lt;code&gt;"YY"&lt;/code&gt;; &lt;code&gt;"MMMM"&lt;/code&gt; full month name, &lt;code&gt;"MMM"&lt;/code&gt; short; &lt;code&gt;"DDDD"&lt;/code&gt; day name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quarter&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;FORMAT( date, "Q" )&lt;/code&gt; (see appendix)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Difference between two dates&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;DATEDIFF( date1, date2, interval )&lt;/code&gt; with interval day, month or year&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Working days between dates&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;NETWORKDAYS&lt;/code&gt; (excludes weekends, optionally holidays)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build a date&lt;/td&gt;
&lt;td&gt;&lt;code&gt;DATE&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week number&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;WEEKNUM&lt;/code&gt;; return type 1 = week starts Sunday, 2 = week starts Monday&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Order Month Name = FORMAT( Orders[Order Date], "MMMM" )
Days to Delivery = DATEDIFF( Orders[Order Date], Orders[Delivery Date], DAY )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s7scuj3xw0z4nqg2sij.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s7scuj3xw0z4nqg2sij.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlw7icl745ou4vpmglxd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlw7icl745ou4vpmglxd.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdgpj1tf7uqxz73acjvob.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdgpj1tf7uqxz73acjvob.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzfs8qu9dvg0teswtgfy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzfs8qu9dvg0teswtgfy.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  7.9 Time intelligence
&lt;/h3&gt;

&lt;p&gt;Time-intelligence functions form a category in the course structure, but no specific functions or examples were covered, so none are presented here as demonstrated content.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Understanding DAX Outputs
&lt;/h2&gt;

&lt;p&gt;Before writing a formula, the first question is: &lt;strong&gt;what output is required?&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Required output&lt;/th&gt;
&lt;th&gt;DAX object&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A single value&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Measure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Appears in the Fields pane; usually shown in a &lt;strong&gt;Card&lt;/strong&gt; visual&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A column in which each row is evaluated separately&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Calculated column&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Adds an entire column to a table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A new table&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Calculated table&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;For example, data for a specific date or specific counties&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A measure is evaluated according to the &lt;strong&gt;filter context&lt;/strong&gt; in which it is used, so it responds to slicers and to the rows and columns of a visual. A calculated column is evaluated row by row when the model is refreshed. Choosing the wrong object type is a common source of errors even when the formula itself is correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. From DAX Functions to Business Questions
&lt;/h2&gt;

&lt;p&gt;DAX is not an exercise in memorising formulas. The essential step is understanding the question and identifying what to filter, what to aggregate and what output is needed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Business question&lt;/th&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total sales for today&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Filter the sum to that date: &lt;code&gt;CALCULATE( SUM( sales ), date = TODAY() )&lt;/code&gt;, or &lt;code&gt;CALCULATE( [Total Sales], date = TODAY() )&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;How many sales&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Count rows or transactions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;How many customers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Counts people who completed transactions, regardless of how many items each bought; &lt;code&gt;COUNTROWS&lt;/code&gt; or &lt;code&gt;COUNT&lt;/code&gt; on the customer ID. Different approaches can give the same answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;How many &lt;em&gt;different&lt;/em&gt; customers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A distinct count (&lt;code&gt;DISTINCTCOUNT&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total yield or market price for a county&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;CALCULATE&lt;/code&gt; with a county filter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Average market price for tomatoes in the dry season&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;CALCULATE&lt;/code&gt; with &lt;code&gt;AVERAGE&lt;/code&gt; and two conditions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These calculations assume tables that are organised sensibly. If the same customer is recorded under several spellings, "how many different customers" cannot be answered correctly. This is the bridge to data modelling.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Data Modelling: Structuring Data for Analysis
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Data modelling&lt;/strong&gt; is the process of organizing tables and defining the relationships between them so that data can be correctly filtered, analyzed and summarized. It involves two tasks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;defining and organizing tables;&lt;/li&gt;
&lt;li&gt;defining relationships between tables.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Why it matters
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Better reporting structure.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reduced repetition.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fewer errors.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ease of creating reports.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A well-designed model also keeps DAX simple, supports query performance (each visual issues a query against the model, and fact/dimension designs suit how those queries filter, group and summarise), improves readability and maintainability, and scales more gracefully as new subject areas are added.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Flat Table
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;flat table&lt;/strong&gt; holds all information in one table. The Kenya crops dataset began this way: county, crop type, planted area, revenue and price all sit in each row.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FLAT TABLE   (illustrative diagram)
┌────────────────────────────────────────────────────────────┐
│ County | Crop  | Farmer | Season | Area | Yield | Revenue … │
│ Kiambu | Maize | F001   | Dry    | 12   | …     | …         │
│ Kiambu | Maize | F002   | Dry    | 8    | …     | …         │
│ Kericho| Tea   | F003   | Wet    | 20   | …     | …         │
└────────────────────────────────────────────────────────────┘
   descriptive text is repeated on every row
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Assessment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Advantages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simple to understand; quick to start; no relationships to configure; adequate for a small, single-topic dataset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Disadvantages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Descriptive values repeat on every row; corrections must be made in many places; tables become very wide; descriptions and measurements are mixed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;When appropriate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Small, one-off analyses with a single subject&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Power BI implications&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;More repeated text is stored; no reusable dimensions to filter from; scales poorly as subjects are added&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The multi-table alternative stores &lt;strong&gt;facts&lt;/strong&gt; (such as transactions or hospital visits) separately from &lt;strong&gt;detail tables&lt;/strong&gt; (such as doctor details), linked to the main table to avoid repetition.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Fact Tables and Dimension Tables
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Fact table
&lt;/h3&gt;

&lt;p&gt;A fact table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keeps the specific records, events or transactions to be analyzed (for example, all patients who visit a hospital);&lt;/li&gt;
&lt;li&gt;Measures something;&lt;/li&gt;
&lt;li&gt;Contains numerical records, metrics and quantities;&lt;/li&gt;
&lt;li&gt;Allows the &lt;strong&gt;grain&lt;/strong&gt; of the data to be understood;&lt;/li&gt;
&lt;li&gt;Contains both primary and foreign keys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Grain (granularity)&lt;/strong&gt; is a description of what each record in a table represents. It indicates &lt;em&gt;what was done&lt;/em&gt;. The test for any fact table is: &lt;strong&gt;what does one row represent?&lt;/strong&gt; In the hospital example, one row of &lt;code&gt;FactVisit&lt;/code&gt; represents one patient visit. If the grain is unclear, counts and totals become unreliable; for instance, if a visit were repeated once per test performed, "number of visits" would be overstated.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Generic examples :&lt;/em&gt; &lt;code&gt;FactSales&lt;/code&gt; (one row per sale line), &lt;code&gt;FactOrders&lt;/code&gt; (one row per order), &lt;code&gt;FactTransactions&lt;/code&gt; (one row per payment).&lt;/p&gt;

&lt;h3&gt;
  
  
  Dimension table
&lt;/h3&gt;

&lt;p&gt;A dimension table&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Describes&lt;/strong&gt; something, giving context for the events held in the fact table (for example, patient details);&lt;/li&gt;
&lt;li&gt;Contains descriptive text content;&lt;/li&gt;
&lt;li&gt;Has a &lt;strong&gt;primary key&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example: &lt;code&gt;DimPatient&lt;/code&gt;, &lt;code&gt;DimDoctor&lt;/code&gt;, &lt;code&gt;DimDepartment&lt;/code&gt;, &lt;code&gt;DimCustomer&lt;/code&gt;, &lt;code&gt;DimProduct&lt;/code&gt;, &lt;code&gt;DimDate&lt;/code&gt;, &lt;code&gt;DimLocation&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Fact table&lt;/th&gt;
&lt;th&gt;Dimension table&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Answers&lt;/td&gt;
&lt;td&gt;What happened? How much?&lt;/td&gt;
&lt;td&gt;Who, what, where, when?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contents&lt;/td&gt;
&lt;td&gt;Numbers, metrics, keys&lt;/td&gt;
&lt;td&gt;Descriptive attributes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One row represents&lt;/td&gt;
&lt;td&gt;One event (the grain)&lt;/td&gt;
&lt;td&gt;One entity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keys&lt;/td&gt;
&lt;td&gt;Primary and foreign keys&lt;/td&gt;
&lt;td&gt;Primary key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical size&lt;/td&gt;
&lt;td&gt;Many rows&lt;/td&gt;
&lt;td&gt;Fewer rows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Separating them means each description is stored once, while the fact table stays focused on events.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. Star Schema
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;star schema&lt;/strong&gt; consists of one central fact table surrounded by dimension tables, each connected to the fact table.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 DimDepartment
                       │
                       │
   (other dimension) ── FactVisit ── DimPatient
                       │
                       │
                   DimDoctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Advantages&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Simple, readable model diagram.&lt;/li&gt;
&lt;li&gt;Descriptions are stored once, reducing repetition.&lt;/li&gt;
&lt;li&gt;Dimensions filter and the fact table is summarised, which matches the way Power BI visuals query a model.&lt;/li&gt;
&lt;li&gt;Simpler DAX: measures aggregate fact columns while slicers come from dimensions.&lt;/li&gt;
&lt;li&gt;Easy report creation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Disadvantages&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Requires planning, since a flat table must be split into facts and dimensions.&lt;/li&gt;
&lt;li&gt;Some descriptive values still repeat within a dimension (for example, a county name for every patient in that county).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Appropriate use:&lt;/strong&gt; most business intelligence and reporting models.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. Snowflake Schema
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;snowflake schema&lt;/strong&gt; works like a star schema, but the dimension tables are broken down into several smaller related tables. The structure is hierarchical: fact → dimension → sub-dimension.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; DimCountry ── DimCounty ── DimPatient ── FactVisit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Star&lt;/th&gt;
&lt;th&gt;Snowflake&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dimensions&lt;/td&gt;
&lt;td&gt;One table per dimension&lt;/td&gt;
&lt;td&gt;Dimension split into related sub-tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redundancy&lt;/td&gt;
&lt;td&gt;Some repetition inside dimensions&lt;/td&gt;
&lt;td&gt;Less repetition (more normalized)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relationship hops from slicer to fact&lt;/td&gt;
&lt;td&gt;One&lt;/td&gt;
&lt;td&gt;Two or more&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model diagram&lt;/td&gt;
&lt;td&gt;Compact&lt;/td&gt;
&lt;td&gt;Larger, more tables&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Advantages:&lt;/strong&gt;&lt;br&gt;
-Less repeated data within dimensions; &lt;br&gt;
-Explicit hierarchies; &lt;br&gt;
-A shared lookup (such as county) can be maintained in one place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disadvantages:&lt;/strong&gt; &lt;br&gt;
-More tables and relationships, &lt;br&gt;
-Longer filter paths, and a &lt;br&gt;
-Busier model that is harder for report builders to navigate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Appropriate situations:&lt;/strong&gt; when a sub-dimension is shared by several dimensions, or when the source data is already normalized. &lt;/p&gt;
&lt;h2&gt;
  
  
  15. Comparing Flat, Star and Snowflake Schemas
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Flat table&lt;/th&gt;
&lt;th&gt;Star schema&lt;/th&gt;
&lt;th&gt;Snowflake schema&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Structure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Everything in one table&lt;/td&gt;
&lt;td&gt;One fact table plus dimension tables&lt;/td&gt;
&lt;td&gt;Fact table plus dimensions split into sub-tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Number of tables&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Few&lt;/td&gt;
&lt;td&gt;Most&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Redundancy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low to moderate&lt;/td&gt;
&lt;td&gt;Lowest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simplest to set up; harder to maintain as it grows&lt;/td&gt;
&lt;td&gt;Moderate; easy to read&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;DAX / reporting&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Adequate for simple cases; filtering limited to that one table&lt;/td&gt;
&lt;td&gt;Simple: measures on facts, slicers from dimensions&lt;/td&gt;
&lt;td&gt;Works, but filters pass through more relationships&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Wide tables with repeated text can be inefficient&lt;/td&gt;
&lt;td&gt;Recommended by Microsoft for Power BI&lt;/td&gt;
&lt;td&gt;Extra relationship hops; usually no benefit unless the structure is needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scalability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Good, with growing complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Maintainability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Corrections repeated across many rows&lt;/td&gt;
&lt;td&gt;Fix a description once&lt;/td&gt;
&lt;td&gt;Good for hierarchies; heavier to manage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Appropriate use&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Small, single-topic analysis&lt;/td&gt;
&lt;td&gt;Most BI projects&lt;/td&gt;
&lt;td&gt;Shared hierarchies; normalised sources&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  16. Primary Keys and Foreign Keys
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;primary key&lt;/strong&gt; is a unique identifier column that gives each record in a table a unique identity.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;foreign key&lt;/strong&gt; is a primary key from one table that is used in a different table.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;common column&lt;/strong&gt; is required before a relationship can be created.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Example.&lt;/strong&gt; &lt;code&gt;CustomerID&lt;/code&gt; is unique in &lt;code&gt;DimCustomer&lt;/code&gt; (one row per customer) but may appear many times in &lt;code&gt;FactSales&lt;/code&gt;, because one customer can make many purchases. &lt;code&gt;DimCustomer[CustomerID]&lt;/code&gt; is therefore a primary key, &lt;code&gt;FactSales[CustomerID]&lt;/code&gt; is a foreign key, and the relationship is &lt;strong&gt;one-to-many&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Illustrative example:&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DimCustomer&lt;/code&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID (PK)&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Brian&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;FactSales&lt;/code&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;SaleID&lt;/th&gt;
&lt;th&gt;CustomerID (FK)&lt;/th&gt;
&lt;th&gt;Amount&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;S1&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S2&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;S3&lt;/td&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;450&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Referential integrity&lt;/strong&gt; means every foreign-key value in the fact table has a matching primary key in the dimension. If &lt;code&gt;FactSales&lt;/code&gt; contained &lt;code&gt;C9&lt;/code&gt; but &lt;code&gt;DimCustomer&lt;/code&gt; had no &lt;code&gt;C9&lt;/code&gt;, those sales would not attach to any customer and would appear under a blank member in visuals.&lt;/p&gt;
&lt;h2&gt;
  
  
  17. Relationships in Power BI
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;relationship&lt;/strong&gt; shows how tables are connected. It is created by matching a primary key in a dimension table with a foreign key in the fact table. Relationships are necessary because, once data is distributed across tables, Power BI needs to know which fact rows belong to which dimension rows. They are what allow &lt;strong&gt;filters to travel&lt;/strong&gt; from one table to another.&lt;/p&gt;

&lt;p&gt;Key properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keys:&lt;/strong&gt; primary key ↔ foreign key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cardinality:&lt;/strong&gt; how rows in one table relate to rows in another (Section 18).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-filter direction:&lt;/strong&gt; single or both (Section 19).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Referential integrity&lt;/strong&gt; (Section 16).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Active and inactive relationships:&lt;/strong&gt; only one relationship between two tables can be active at a time. Additional relationships can exist but are inactive (shown as dashed lines in Model view) and do not filter unless a DAX measure activates them with &lt;code&gt;USERELATIONSHIP&lt;/code&gt;. A relationship must exist and be active for columns from different tables to be used together.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  18. Relationship Cardinality
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cardinality&lt;/strong&gt; describes how data in one table relates to data in another, based on the count of matching rows. It should be distinguished from &lt;em&gt;column cardinality&lt;/em&gt;, which refers to how unique the values in a single column are.&lt;/p&gt;
&lt;h3&gt;
  
  
  One-to-many (1:*), the most common type
&lt;/h3&gt;

&lt;p&gt;A single row in one table matches multiple rows in the other.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Example: one patient (dimension) → many visits (fact); &lt;em&gt;illustrative:&lt;/em&gt; one &lt;code&gt;DimCustomer&lt;/code&gt; row → many &lt;code&gt;FactSales&lt;/code&gt; rows.&lt;/li&gt;
&lt;li&gt;Common in dimensional models because every dimension-to-fact link is of this type: the "one" side holds unique keys and the "many" side holds the events.
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DimCustomer (1) ─────────────&amp;lt; FactSales (*)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;em&gt;Many-to-one&lt;/em&gt; is the same relationship read from the other direction.&lt;/p&gt;
&lt;h3&gt;
  
  
  One-to-one (1:1)
&lt;/h3&gt;

&lt;p&gt;One row in a table links to only one row in another table. It arises when there is no separate information that justifies splitting the table, and it relates to &lt;strong&gt;normalisation&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Illustrative example:&lt;/em&gt; &lt;code&gt;Employee&lt;/code&gt; and &lt;code&gt;EmployeeBadge&lt;/code&gt;, with exactly one badge record per employee.&lt;/li&gt;
&lt;li&gt;Before using it, consider whether the two tables should simply be combined into one (for example by a merge in Power Query). Filters flow in both directions in a 1:1 relationship.
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Employee (1) ─────────── (1) EmployeeBadge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Many-to-many (:)
&lt;/h3&gt;

&lt;p&gt;Multiple rows in one table link to multiple rows in another.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Illustrative example:&lt;/em&gt; students and courses.&lt;/li&gt;
&lt;li&gt;It requires careful modelling because neither side has unique keys, so filters can match many rows on both sides and totals can become hard to interpret. A &lt;strong&gt;bridge (junction) table&lt;/strong&gt; converts it into two one-to-many relationships.
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Students (*) ───────────── (*) Courses
   safer:  Students (1)──&amp;lt;  Enrolments  &amp;gt;──(1) Courses
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzdk8tmbm9nqwonwwcwdi.png" alt=" " width="800" height="450"&gt;
&lt;/h2&gt;
&lt;h2&gt;
  
  
  19. Filter Direction
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Filter direction&lt;/strong&gt; controls which table filters which when a user interacts with a slicer or visual.&lt;/p&gt;
&lt;h3&gt;
  
  
  Single-direction filtering
&lt;/h3&gt;

&lt;p&gt;Filters flow from the "one" side (dimension) to the "many" side (fact).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DimProduct
    │
    ▼
FactSales
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Selecting a product in a &lt;code&gt;DimProduct&lt;/code&gt; slicer filters the corresponding rows in &lt;code&gt;FactSales&lt;/code&gt;, so a revenue card then shows only that product's revenue (&lt;em&gt;illustrative&lt;/em&gt;). This is the default for a one-to-many relationship and the preferred setting in a star schema.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bidirectional filtering
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DimProduct
    ▲
    ▼
FactSales
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Filters flow in both directions. It should be used sparingly because it can cause:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ambiguous filter paths&lt;/strong&gt;, where multiple routes between two tables leave it unclear which one Power BI should use, potentially deactivating a relationship or producing route-dependent results;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;unnecessary complexity&lt;/strong&gt;, since each bidirectional relationship makes the model harder to reason about;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;unexpected filtering behaviour&lt;/strong&gt;, as slicers begin filtering one another and numbers change in ways report users do not expect;&lt;/li&gt;
&lt;li&gt;reduced performance on large models.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Guideline:&lt;/strong&gt; keep relationships single-direction from dimension to fact, and use &lt;em&gt;Both&lt;/em&gt; only for a specific, understood reason.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgrqx0lm0298ye3o3ktk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgrqx0lm0298ye3o3ktk.png" alt=" " width="800" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  20. Establishing Relationships in Power BI
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Identify the matching keys:&lt;/strong&gt; the primary key in the dimension and the foreign key in the fact table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create the relationship&lt;/strong&gt; by either:

&lt;ul&gt;
&lt;li&gt;drag and drop, from primary key to foreign key in Model view; or&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modeling → Manage relationships&lt;/strong&gt; on the ribbon.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set the cardinality&lt;/strong&gt; and the &lt;strong&gt;cross-filter direction&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Preparing keys when none exist
&lt;/h3&gt;

&lt;p&gt;When a flat table is split into dimensions, the following workflow applies:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Clean the data.&lt;/li&gt;
&lt;li&gt;Create dimension tables, working out which columns belong to each dimension.&lt;/li&gt;
&lt;li&gt;Give each dimension a &lt;strong&gt;unique identifier&lt;/strong&gt;. If none exists, create one: right-click on the table → &lt;strong&gt;Add index column (from 1)&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk4e1tt2joghgrwr2xh2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk4e1tt2joghgrwr2xh2.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In the original table, find and replace the descriptive name with the unique index, which then serves as the &lt;strong&gt;foreign key&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Clean the data again and check for errors.&lt;/li&gt;
&lt;li&gt;Create relationships and ensure they are active.&lt;/li&gt;
&lt;li&gt;Calculate (DAX).&lt;/li&gt;
&lt;li&gt;Display (dashboard).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The same result can also be achieved by merging the original table with the new dimension in Power Query, which leads to the next topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  21. Joins in Power Query
&lt;/h2&gt;

&lt;p&gt;Relationships connect tables inside the model. A &lt;strong&gt;join&lt;/strong&gt; combines tables earlier, during data preparation. A join combines two tables using matching columns identified through keys, and is performed in Power Query because the table is being transformed to create a new, merged table. &lt;br&gt;
In Power Query the command is &lt;strong&gt;Merge Queries&lt;/strong&gt;. Which table's rows are kept in full depends on the join type.&lt;/p&gt;
&lt;h3&gt;
  
  
  Example used for all six joins
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Illustrative example&lt;/em&gt; &lt;br&gt;
Join key: &lt;code&gt;CustomerID&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Customers (left table)&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Brian&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C3&lt;/td&gt;
&lt;td&gt;Chebet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Orders (right table)&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;O1&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;O2&lt;/td&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;O3&lt;/td&gt;
&lt;td&gt;C4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Brian (C2) and Chebet (C3) have no orders, and order O3 belongs to C4, who is not in the Customers table.&lt;/p&gt;
&lt;h3&gt;
  
  
  21.1 Inner join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;only matching records&lt;/strong&gt; in both tables. Use when only records existing on both sides are wanted.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customers ∩ Orders         only the overlap
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F454u6098e0lggtfnmznl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F454u6098e0lggtfnmznl.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2htsam2fi0lvs9s3w29p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2htsam2fi0lvs9s3w29p.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  21.2 Left outer join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;all rows from the left (first) table&lt;/strong&gt; and only the matching rows from the right. Use to list all customers whether or not they ordered.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Brian&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C3&lt;/td&gt;
&lt;td&gt;Chebet&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;all of Customers + matching Orders
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpbkz78ns7lrpkvdv58mb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpbkz78ns7lrpkvdv58mb.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiyvbuznl2pcdydvqfmnm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiyvbuznl2pcdydvqfmnm.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  21.3 Right outer join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;all rows from the right (second) table&lt;/strong&gt; and the matching rows from the left.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C4&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;O3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;matching Customers + all of Orders
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0pns2gm3wee8apo9nzm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0pns2gm3wee8apo9nzm.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  21.4 Full outer join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;all records from both tables&lt;/strong&gt;, whether or not they match.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C1&lt;/td&gt;
&lt;td&gt;Amina&lt;/td&gt;
&lt;td&gt;O2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Brian&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C3&lt;/td&gt;
&lt;td&gt;Chebet&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C4&lt;/td&gt;
&lt;td&gt;&lt;em&gt;null&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;O3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customers ∪ Orders         everything
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0nwzu6v68rhzvr5kypzd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0nwzu6v68rhzvr5kypzd.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  21.5 Left anti join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;only the rows in the left table that have no match&lt;/strong&gt; in the right. Useful for finding customers who never ordered, and for data-quality checks.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;th&gt;CustomerName&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;C2&lt;/td&gt;
&lt;td&gt;Brian&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C3&lt;/td&gt;
&lt;td&gt;Chebet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customers − Orders         left only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffxm13tswd1wackd0ccsx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffxm13tswd1wackd0ccsx.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6hv0y7heydozmf0hyo0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6hv0y7heydozmf0hyo0.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  21.6 Right anti join
&lt;/h3&gt;

&lt;p&gt;Keeps &lt;strong&gt;only the rows in the right table that have no match&lt;/strong&gt; in the left. Useful for finding orders whose customer is missing, which is a referential-integrity problem (Section 16).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OrderID&lt;/th&gt;
&lt;th&gt;CustomerID&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;O3&lt;/td&gt;
&lt;td&gt;C4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Orders − Customers         right only
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhh4q8zebnbg7fx4vdbni.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhh4q8zebnbg7fx4vdbni.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Summary
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Join&lt;/th&gt;
&lt;th&gt;Rows kept&lt;/th&gt;
&lt;th&gt;Result in the example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inner&lt;/td&gt;
&lt;td&gt;Matches only&lt;/td&gt;
&lt;td&gt;2 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Left outer&lt;/td&gt;
&lt;td&gt;All left + matching right&lt;/td&gt;
&lt;td&gt;4 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Right outer&lt;/td&gt;
&lt;td&gt;All right + matching left&lt;/td&gt;
&lt;td&gt;3 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full outer&lt;/td&gt;
&lt;td&gt;Everything&lt;/td&gt;
&lt;td&gt;5 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Left anti&lt;/td&gt;
&lt;td&gt;Left rows with no match&lt;/td&gt;
&lt;td&gt;2 rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Right anti&lt;/td&gt;
&lt;td&gt;Right rows with no match&lt;/td&gt;
&lt;td&gt;1 row&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;


&lt;h2&gt;
  
  
  22. Join Output Illustrations
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; Left table = L,  Right table = R,  overlap = matched rows

 Inner       :        [ L ∩ R ]
 Left outer  :   [ L ........ ∩ ]       all of L, matched R
 Right outer :        [ ∩ ........ R ]   matched L, all of R
 Full outer  :   [ L ........ ∩ ........ R ]
 Left anti   :   [ L only ]
 Right anti  :                [ R only ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  23. Power Query Joins vs Power BI Relationships
&lt;/h2&gt;

&lt;p&gt;Both mechanisms connect tables through matching keys, but they occur at different stages and do different things.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; Power Query                         Data model
 ───────────                         ──────────
 Merge / Join  ──►  Data           ──►  Relationships  ──►  DAX / Visuals
 (combine rows)     transformation       (connect tables)
 "preparation"                           "modelling"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Does a Power Query merge physically combine data?&lt;/strong&gt; Yes. A merge creates a new table, or adds columns to an existing one, in which matched rows sit side by side. After &lt;em&gt;Close &amp;amp; Apply&lt;/em&gt;, that wider table is what is loaded into the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a relationship physically combine tables?&lt;/strong&gt; No. A relationship creates no new table and moves no data. Both tables remain separate. It is a rule in the model stating that when one table is filtered, the other is filtered through a shared key. Power BI applies it at calculation time, when a visual or measure needs columns from both tables.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When does each happen?&lt;/strong&gt; Joins happen during preparation (Power Query); relationships are defined during modelling (Model view), after loading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When is a merge preferable to a relationship?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;To add a column from another table, such as bringing a county name into a table that only held a county code (&lt;em&gt;illustrative&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;To flatten a snowflake by merging sub-dimensions into one dimension.&lt;/li&gt;
&lt;li&gt;To filter rows by existence in another table, using anti joins.&lt;/li&gt;
&lt;li&gt;To bring a key into a fact table when creating dimensions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A relationship is preferable when tables hold different kinds of information (dimension and fact) and should stay separate while still filtering one another.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What can excessive merging cause?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wider tables with more columns.&lt;/li&gt;
&lt;li&gt;Greater redundancy, since descriptions repeat on every row (the flat-table problem).&lt;/li&gt;
&lt;li&gt;Loss of dimensional structure, because facts and dimensions blend into one table and clean, reusable dimensions disappear.&lt;/li&gt;
&lt;li&gt;Higher model complexity and maintenance cost, and row multiplication if the key is not unique on one side of the merge.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why do fact and dimension tables stay separate?&lt;/strong&gt; A star schema works because dimensions filter and facts summarise. Merging everything into one table abandons that structure. Keeping the tables separate and relating them stores each description once, keeps the fact table focused on its grain, and keeps DAX simple.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Power Query merge (join)&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Power BI relationship&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Stage&lt;/td&gt;
&lt;td&gt;Data preparation&lt;/td&gt;
&lt;td&gt;Data modelling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Location&lt;/td&gt;
&lt;td&gt;Power Query Editor&lt;/td&gt;
&lt;td&gt;Model view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Physically combines data?&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes&lt;/strong&gt;: creates a merged table&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No&lt;/strong&gt;: tables stay separate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;A new table or extra columns&lt;/td&gt;
&lt;td&gt;A filter path between tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Controlled by&lt;/td&gt;
&lt;td&gt;Join kind (six types)&lt;/td&gt;
&lt;td&gt;Cardinality and filter direction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When applied&lt;/td&gt;
&lt;td&gt;At refresh / load&lt;/td&gt;
&lt;td&gt;At query time, when visuals and measures run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Cleaning, enriching, flattening, checking&lt;/td&gt;
&lt;td&gt;Linking facts to dimensions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;In short: join during preparation; relate during modelling.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  24. From Model to Visualisation
&lt;/h2&gt;

&lt;p&gt;Once data has been cleaned, modelled, related and calculated, reports can be built. In &lt;strong&gt;Report view&lt;/strong&gt;, a chart type is selected and the fields to plot are added. Visuals are interactive: selecting part of one visual filters the others, including KPIs, and even the components of a bar are interactive.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Visual&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bar chart&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Comparing categories&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pie / Donut / Treemap&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Part-to-whole relationships; treemaps suit many categories&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scatter&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Relationship between two numeric variables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Card&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Displaying a KPI or single headline value (typically a measure)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Map&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Geographic data (bubble, filled and shape maps)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Gauge / KPI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Targets and comparison&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Slicer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Filtering by user selection&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These visuals depend on the earlier stages. The card shows a &lt;code&gt;CALCULATE&lt;/code&gt; measure (Section 7.4); the slicer filters through a relationship (Sections 17 and 19); chart categories come from clean dimension columns (Section 5); and the &lt;em&gt;Total Revenue all county&lt;/em&gt; card (Section 7.5) shows how &lt;code&gt;ALL&lt;/code&gt; makes a KPI ignore slicers deliberately.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc09wpbhiqy2t6i3kek9w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc09wpbhiqy2t6i3kek9w.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  25. Publishing to Power BI Service
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Key terms
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;.PBIX:&lt;/strong&gt; the Power BI Desktop file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic model:&lt;/strong&gt; the entire data ecosystem: tables, relationships and connections, measures, and the model itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Report:&lt;/strong&gt; the visuals, graphs and dashboards that report users read to evaluate the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workspace:&lt;/strong&gt; a collaborative area in Power BI Service where a team creates, manages and stores Power BI content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Power BI Service:&lt;/strong&gt; the browser version of Power BI, used to create, share and consume content.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What publishing does
&lt;/h3&gt;

&lt;p&gt;Publishing sends the report and the semantic model from the &lt;code&gt;.pbix&lt;/code&gt; file to the selected workspace.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Power BI Desktop (.pbix)
          │
          ▼  Home → Publish → select destination
Power BI Service
          │
          ▼
      Workspace
          │
          ▼
Semantic model  +  Report
          │
          ▼
Sharing / Collaboration  (workspace roles and security)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The semantic model is the same model built in Model view, now hosted in the cloud, so earlier modelling decisions affect every report built on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Steps
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Before publishing:&lt;/strong&gt; sign in to Power BI from the Desktop app, then sign in to Power BI Service online.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a workspace:&lt;/strong&gt; open the left pane → &lt;strong&gt;Workspaces&lt;/strong&gt; and complete the details (name, description, image, domain, Power BI Pro where external users need access, contact list).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish:&lt;/strong&gt; in Desktop, &lt;strong&gt;Home → Publish → select destination&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open the workspace and refresh&lt;/strong&gt; to see the report and semantic model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share:&lt;/strong&gt; subscriptions allow a particular report page to be shared with an intended person; a report with multiple pages can also be shared as a link.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9vn5fo0cl0eqr05fhyi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9vn5fo0cl0eqr05fhyi.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cj2ykvs6d7bp25q2ugy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cj2ykvs6d7bp25q2ugy.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  26. Workspace Roles and Security
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Workspace roles
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Admin&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Controls the workspace and access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Member&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Can collaborate and publish into the workspace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Contributor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Can create and change content, with limited access functions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Viewer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Can consume content without editing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Exact permission lists should be confirmed against current Microsoft documentation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Row-level security (RLS)
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Row-level security&lt;/strong&gt; controls which rows of data a user can see based on their assigned role, to keep data secure. Rules specify who needs access to what, and roles can be created before the report is shared and tested with &lt;strong&gt;View as&lt;/strong&gt;.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Define&lt;/strong&gt; roles and their DAX filter rules in Power BI Desktop (&lt;strong&gt;Modeling → Manage roles&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish&lt;/strong&gt;, then &lt;strong&gt;assign users or groups&lt;/strong&gt; to each role in Power BI Service. A defined role is not enforced until users are assigned.&lt;/li&gt;
&lt;li&gt;RLS applies to &lt;strong&gt;Viewers&lt;/strong&gt;. Members of the &lt;strong&gt;Admin, Member and Contributor&lt;/strong&gt; roles have edit permission on the semantic model, so RLS does not restrict them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test&lt;/strong&gt; with &lt;em&gt;View as&lt;/em&gt; before sharing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;RLS filters travel across relationships like any other filter and by default follow single direction, which is a further reason to keep filter directions simple.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqsaj93myot2xzrlclqo7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqsaj93myot2xzrlclqo7.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  27. Recommended Power BI Model
&lt;/h2&gt;

&lt;p&gt;For a typical business intelligence project, a &lt;strong&gt;star schema&lt;/strong&gt; is the recommended design. The reasoning is that it fits how Power BI queries models and the types of questions BI projects ask, not that it is universally superior.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Flat table&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Star schema&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;Snowflake&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Query / report performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Repeated text bloats wide tables&lt;/td&gt;
&lt;td&gt;Matches how Power BI visuals query the model; recommended by Microsoft&lt;/td&gt;
&lt;td&gt;Extra hops between tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;DAX simplicity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Simple at first; awkward as questions grow&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Simple:&lt;/strong&gt; measures on facts, slicers from dimensions&lt;/td&gt;
&lt;td&gt;More relationships to reason about&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Readability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hard to tell what each column represents&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Clear:&lt;/strong&gt; facts in the centre, dimensions around&lt;/td&gt;
&lt;td&gt;Larger diagram&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scalability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Good:&lt;/strong&gt; add a dimension or fact&lt;/td&gt;
&lt;td&gt;Good, with added complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Data redundancy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Lowest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Maintainability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Same value fixed in many rows&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fix a description once&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Good for hierarchies; heavier to manage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ease of report creation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One place for everything, but limited filtering&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Easy:&lt;/strong&gt; dimension fields for slicers and axes, fact fields for values&lt;/td&gt;
&lt;td&gt;Builders must understand the chain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Filter propagation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Nothing to manage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Predictable:&lt;/strong&gt; dimension → fact&lt;/td&gt;
&lt;td&gt;Passes through several tables&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Lowest initially&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Moderate&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  When another design is appropriate
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flat table:&lt;/strong&gt; a small, one-off analysis with one subject, where building dimensions costs more than it returns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Snowflake:&lt;/strong&gt; when a sub-dimension is genuinely shared (for example, a county table used by both patients and clinics) or the source is already normalised and flattening is not worthwhile. Merging the sub-tables in Power Query to obtain a simpler star should be considered first.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Recommended relationship design
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One fact table&lt;/strong&gt; at the centre, with a grain that can be stated in one sentence (for example, &lt;em&gt;one row per visit&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One dimension table per descriptive subject&lt;/strong&gt; (patient, doctor, department, date, location), each with a &lt;strong&gt;unique primary key&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Foreign keys in the fact table&lt;/strong&gt; matching those primary keys. Where no natural key exists, create an &lt;strong&gt;index column&lt;/strong&gt; and carry it into the fact table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cardinality:&lt;/strong&gt; one-to-many (1:*) from each dimension to the fact table. Avoid many-to-many unless unavoidable, using a bridge table where needed, and consider merging one-to-one pairs into a single table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filter direction:&lt;/strong&gt; single, from dimension to fact. Use &lt;em&gt;Both&lt;/em&gt; only for a specific, understood reason, to avoid ambiguous paths.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Active relationships&lt;/strong&gt; on the main path; an inactive relationship only where a measure deliberately uses &lt;code&gt;USERELATIONSHIP&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Referential integrity checked in Power Query&lt;/strong&gt; using a right anti join (fact keys with no dimension match) before loading.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; DimDepartment   DimDoctor
        \          /
         \        /
 DimPatient ── FactVisit ── DimDate
   (1)──────────►(*)         (1)──►(*)      single-direction, 1:* from every dimension
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;(Illustrative design based on the hospital example, with a date dimension added.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fislk3pctockjfocbj02n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fislk3pctockjfocbj02n.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
F&lt;/p&gt;

&lt;h2&gt;
  
  
  28. The Complete Workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; 1. Understand the problem / question
              ↓
 2. Identify the data source
              ↓
 3. Import / connect data
              ↓
 4. Inspect the raw data
              ↓
 5. Clean and transform with Power Query
              ↓
 6. Understand the structure of the data
              ↓
 7. Build the data model
              ↓
 8. Create fact and dimension tables
              ↓
 9. Establish relationships
              ↓
10. Apply appropriate cardinality / filter direction
              ↓
11. Use DAX to calculate and analyse
              ↓
12. Create visualisations
              ↓
13. Build the report
              ↓
14. Publish to Power BI Service
              ↓
15. Share, secure and collaborate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;How the stages connect&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Steps 1–3:&lt;/strong&gt; The business question determines which data is needed and where it lives. Import brings a copy into Desktop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps 4–5:&lt;/strong&gt; Imported data is rarely analysis-ready. Data types, blanks, errors and duplicates are resolved in Power Query, with Applied Steps recording each action. Joins are available here for combining or checking tables.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps 6–8:&lt;/strong&gt; Clean data is organised: measurements into a fact table with a clear grain, descriptions into dimension tables, usually in a star schema.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps 9–10:&lt;/strong&gt; Primary keys in dimensions link to foreign keys in the fact table through one-to-many, single-direction relationships. The tables then work as one model without being merged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step 11:&lt;/strong&gt; DAX operates on top of the model. &lt;code&gt;CALCULATE&lt;/code&gt;, &lt;code&gt;FILTER&lt;/code&gt; and &lt;code&gt;ALL&lt;/code&gt; depend on filters moving through relationships, so the better the model, the simpler the DAX.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps 12–13:&lt;/strong&gt; Visuals present measures and fields and respond to the filters defined in the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steps 14–15:&lt;/strong&gt; The &lt;code&gt;.pbix&lt;/code&gt; is published to a workspace as a semantic model and report, with roles and RLS governing who can do and see what.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  29. Conclusion
&lt;/h2&gt;

&lt;p&gt;Power BI is best understood as a chain of dependent practices rather than a set of separate features. Data enters through connections to sources such as Excel, CSV, SQL and the web. Its quality must be established in Power Query before it can be trusted. DAX then turns clean data into answers, but its effectiveness depends on a well-structured model: fact tables with a defined grain, dimension tables with unique keys, one-to-many relationships and carefully chosen filter directions. Joins and relationships are complementary tools applied at different stages: joins combine data while it is being prepared, and relationships connect tables once they are loaded. Reports, dashboards and publishing sit at the end of the chain, so their reliability reflects the care taken in every earlier stage.&lt;/p&gt;

&lt;p&gt;For most business intelligence projects, a star schema built on clean, correctly typed data offers the best balance of performance, simplicity, readability and maintainability.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>data</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Extract Job Listings from Indeed — Python Tutorial 2026</title>
      <dc:creator>RAZIX DEVIL NEMESIS (Loki)</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:15:48 +0000</pubDate>
      <link>https://dev.to/razix_devilnemesisloki/extract-job-listings-from-indeed-python-tutorial-2026-1733</link>
      <guid>https://dev.to/razix_devilnemesisloki/extract-job-listings-from-indeed-python-tutorial-2026-1733</guid>
      <description>&lt;h1&gt;
  
  
  Extract Job Listings from Indeed — Python Tutorial 2026
&lt;/h1&gt;

&lt;p&gt;Indeed is the world's largest job search engine with 250M+ monthly visitors. Here's how to extract job listing data programmatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Scrape Indeed?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Market salary analysis&lt;/li&gt;
&lt;li&gt;Job demand trends&lt;/li&gt;
&lt;li&gt;Recruitment automation&lt;/li&gt;
&lt;li&gt;Competitive intelligence&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Quick Start with Apify
&lt;/h2&gt;

&lt;p&gt;Use the &lt;a href="https://apify.com/store" rel="noopener noreferrer"&gt;Indeed Job Scraper&lt;/a&gt; on Apify — no coding needed. Just enter keywords and location, get structured JSON back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Python Method
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;bs4&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BeautifulSoup&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search_indeed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;location&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://www.indeed.com/jobs?q=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;keyword&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;&amp;amp;l=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;location&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;User-Agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Mozilla/5.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="n"&gt;soup&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;BeautifulSoup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lxml&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;jobs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;card&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;soup&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.job_seen_beacon&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;card&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.jobTitle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;company&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;card&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.companyName&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;company&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;jobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;strip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;company&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;strip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;jobs&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Data You Can Extract
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Job title and description&lt;/li&gt;
&lt;li&gt;Company name and rating&lt;/li&gt;
&lt;li&gt;Salary range&lt;/li&gt;
&lt;li&gt;Location&lt;/li&gt;
&lt;li&gt;Posting date&lt;/li&gt;
&lt;li&gt;Application link&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best Practices
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Use polite delays between requests&lt;/li&gt;
&lt;li&gt;Rotate user agents&lt;/li&gt;
&lt;li&gt;Consider using a proxy service&lt;/li&gt;
&lt;li&gt;Check Indeed's terms of service&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Built by an AI agent earning passively. Deploy your own: &lt;a href="https://github.com/JoseLuis6934/omnincome-agent" rel="noopener noreferrer"&gt;Omnincome Agent&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>jobs</category>
      <category>python</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Why Your PDF Text Extraction Returns an Empty String</title>
      <dc:creator>Nikolas Dimitroulakis</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:13:36 +0000</pubDate>
      <link>https://dev.to/nikolas_dimitroulakis_d23/why-your-pdf-text-extraction-returns-an-empty-string-2oli</link>
      <guid>https://dev.to/nikolas_dimitroulakis_d23/why-your-pdf-text-extraction-returns-an-empty-string-2oli</guid>
      <description>&lt;p&gt;You wrote ten lines of Python, pointed it at a PDF, and got back &lt;code&gt;""&lt;/code&gt;. No exception, no warning, nothing.&lt;/p&gt;

&lt;p&gt;The reason is almost always the same: &lt;strong&gt;that PDF has no text layer.&lt;/strong&gt; It is a picture of a document, not a document. Text-layer libraries return nothing and report no error, which makes this one of the quieter bugs you can ship.&lt;/p&gt;

&lt;p&gt;Here is how to detect it and route around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Check for a text layer before you extract
&lt;/h2&gt;

&lt;p&gt;From the shell, &lt;a href="https://poppler.freedesktop.org/" rel="noopener noreferrer"&gt;poppler-utils&lt;/a&gt; answers this in one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pdffonts document.pdf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fonts listed means there is a text layer. Empty output means it is a scan.&lt;/p&gt;

&lt;p&gt;In Python, with &lt;a href="https://pymupdf.readthedocs.io/" rel="noopener noreferrer"&gt;PyMuPDF&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;fitz&lt;/span&gt;  &lt;span class="c1"&gt;# PyMuPDF
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;has_text_layer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;min_chars&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pages_to_check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;fitz&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;pages_to_check&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_text&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;min_chars&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details worth keeping:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check several pages, not just the first.&lt;/strong&gt; Scanned documents often carry a generated cover page with real text in front of forty pages of images. A page-one check sends the whole file down the wrong branch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use a character threshold, not a truthiness check.&lt;/strong&gt; Plenty of scans carry a few stray characters from a header or a watermark. That is enough to pass &lt;code&gt;if text:&lt;/code&gt; and nothing more.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Text-layer PDFs
&lt;/h2&gt;

&lt;p&gt;PyMuPDF is the fastest general-purpose option:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;extract_text_layer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;fitz&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_text&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If columns come out interleaved, &lt;code&gt;pdftotext -layout document.pdf document.txt&lt;/code&gt; often does better, and &lt;a href="https://github.com/jsvine/pdfplumber" rel="noopener noreferrer"&gt;pdfplumber&lt;/a&gt; gives you word-level coordinates when you need to reconstruct the layout yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Scanned PDFs
&lt;/h2&gt;

&lt;p&gt;Scans need OCR. &lt;a href="https://ocrmypdf.readthedocs.io/" rel="noopener noreferrer"&gt;OCRmyPDF&lt;/a&gt; wraps &lt;a href="https://github.com/tesseract-ocr/tesseract" rel="noopener noreferrer"&gt;Tesseract&lt;/a&gt; and writes a text layer back into the file, which means the rest of your pipeline stays unchanged:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ocr_pdf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eng&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ocrmypdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-l&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--skip-text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;extract_text_layer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;-l&lt;/code&gt; to the right language. The default is English, and accuracy drops sharply on anything else.&lt;/p&gt;

&lt;p&gt;For documents the local stack handles badly (poor scans, forms, handwriting), a hosted &lt;a href="https://apyhub.com/apyhub/service/extract-read-data" rel="noopener noreferrer"&gt;OCR document data extraction API&lt;/a&gt; returns layout-aware parsing without a Tesseract install, and an &lt;a href="https://apyhub.com/apyhub/service/extract-text-from-image" rel="noopener noreferrer"&gt;image OCR API&lt;/a&gt; covers loose JPEGs and PNGs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Route per file
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;extract&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ocr_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ocr.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;has_text_layer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;method&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text-layer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;extract_text_layer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;method&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ocr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;ocr_pdf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ocr_output&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Return the method alongside the text and log it. The day output looks wrong, you will know in one query whether the problem was your extraction or the source document.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calling it as a hosted API
&lt;/h2&gt;

&lt;p&gt;If you would rather not run any of this, the &lt;a href="https://apyhub.com/apyhub/service/extract-text-from-pdf" rel="noopener noreferrer"&gt;Extract Text from PDF API&lt;/a&gt; takes a remote URL or a file upload and returns the text in a single &lt;code&gt;data&lt;/code&gt; field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.eu.apyhub.com/extract/text/pdf-url"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"apy-token: YOUR_API_KEY"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"url": "https://example.com/report.pdf"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also accepts a page range and coordinate bounds, which is useful when you only need one section of a long report.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which method for which job
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Good at&lt;/th&gt;
&lt;th&gt;Trade-off&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pdftotext -layout&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;One-off text-layer extraction, keeps columns roughly intact&lt;/td&gt;
&lt;td&gt;No structure, little control over output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PyMuPDF&lt;/td&gt;
&lt;td&gt;Fast bulk extraction with coordinates&lt;/td&gt;
&lt;td&gt;Tables lose row structure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pdfplumber&lt;/td&gt;
&lt;td&gt;Word positions, lined tables&lt;/td&gt;
&lt;td&gt;Noticeably slower on large files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ocrmypdf + Tesseract&lt;/td&gt;
&lt;td&gt;Clean scans, free, runs locally&lt;/td&gt;
&lt;td&gt;Accuracy falls off below 300 DPI and on skewed pages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud OCR (Azure, Google, AWS)&lt;/td&gt;
&lt;td&gt;Poor scans, forms, handwriting&lt;/td&gt;
&lt;td&gt;Per-page cost, data leaves your infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apyhub.com/apyhub/service/extract-text-from-pdf" rel="noopener noreferrer"&gt;Extract Text from PDF API&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Text-layer PDFs by URL or upload, page ranges and coordinate bounds, one call, nothing to install&lt;/td&gt;
&lt;td&gt;50 atoms per call&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  A note on tables
&lt;/h2&gt;

&lt;p&gt;Tables are their own problem. Lined tables (visible borders) extract reasonably well with pdfplumber or &lt;a href="https://camelot-py.readthedocs.io/" rel="noopener noreferrer"&gt;camelot&lt;/a&gt;. Unlined tables usually do not, and OCR makes it worse by returning text in reading order and discarding which cell each value came from.&lt;/p&gt;

&lt;p&gt;When row and column structure matters, a &lt;a href="https://apyhub.com/apyhub/service/extract-table-data" rel="noopener noreferrer"&gt;table extraction API&lt;/a&gt; that preserves layout saves more time than tuning a parser will.&lt;/p&gt;

&lt;h2&gt;
  
  
  The summary
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Detect the text layer per file, never per batch.&lt;/li&gt;
&lt;li&gt;Text layer goes to PyMuPDF.&lt;/li&gt;
&lt;li&gt;No text layer goes to OCR.&lt;/li&gt;
&lt;li&gt;Log which branch each file took.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most of the pain in PDF extraction comes from assuming one method handles everything. Once the routing is in place, each method only has to be good at the job it is actually given.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Every API above is also available through &lt;a href="https://apyhub.com/mcp" rel="noopener noreferrer"&gt;ApyHub MCP&lt;/a&gt;, so an AI agent can call them directly without a hand-written wrapper. You can try any of them from its page in the &lt;a href="https://apyhub.com/catalog" rel="noopener noreferrer"&gt;ApyHub catalog&lt;/a&gt; or in &lt;a href="https://voiden.md/" rel="noopener noreferrer"&gt;Voiden&lt;/a&gt;, the open-source API client, before writing any code.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>pdf</category>
      <category>ocr</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
