<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: confident_prep</title>
    <description>The latest articles on DEV Community by confident_prep (@confident_prep).</description>
    <link>https://dev.to/confident_prep</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4012191%2F281ccd97-d952-44e6-bf08-ac37e33fc1b3.png</url>
      <title>DEV Community: confident_prep</title>
      <link>https://dev.to/confident_prep</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/confident_prep"/>
    <language>en</language>
    <item>
      <title>AWS AIF-C01: Bias, Variance &amp; the AWS Tools That Detect Them</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:13:20 +0000</pubDate>
      <link>https://dev.to/confident_prep/aws-aif-c01-bias-variance-the-aws-tools-that-detect-them-4a6h</link>
      <guid>https://dev.to/confident_prep/aws-aif-c01-bias-variance-the-aws-tools-that-detect-them-4a6h</guid>
      <description>&lt;p&gt;A model can perform well overall and still fail badly for a particular group. That is the central idea behind bias and variance questions on the AWS Certified AI Practitioner exam: training-versus-unseen performance reveals underfitting or overfitting, while differences between demographic groups reveal societal bias. To choose the right detection method, first identify which problem the evidence describes, then match the tool to when the check occurs—point-in-time, continuously after deployment, or on an individual prediction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bias has two meanings, and the scenario tells you which one applies
&lt;/h2&gt;

&lt;p&gt;On the exam, “bias” can refer to either statistical bias or societal bias. These are separate problems, even though both can cause inaccurate predictions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Statistical bias&lt;/strong&gt; means a model is too simple to capture the underlying pattern. It performs poorly on the data used to train it and on data it has never seen. This failure is called &lt;strong&gt;underfitting&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Societal bias&lt;/strong&gt; means outcomes differ systematically across demographic groups. A model might perform well in aggregate while producing much worse results for one group. This is a fairness failure, not necessarily a fitting failure.&lt;/p&gt;

&lt;p&gt;The surrounding language tells you which meaning the question intends:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;References to training accuracy, unseen-data accuracy, overfitting, or underfitting point to statistical bias.&lt;/li&gt;
&lt;li&gt;References to fairness, demographic groups, or unequal outcomes point to societal bias.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Consider an illustrative example. A model scores poorly on both its training set and its test set. That is high statistical bias: the model never captured the pattern. Now suppose another model scores well on both sets overall but performs much worse for rural applicants. That is societal bias. The second model may generalize well in aggregate and still be unfair.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7zu3c8p2blzq503kl2u1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7zu3c8p2blzq503kl2u1.png" alt="Training and unseen accuracy indicate statistical bias; group outcomes indicate societal bias." width="800" height="195"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The same word points to different problems depending on the evidence in the scenario.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This distinction prevents a common exam mistake: calling every group disparity “overfitting.” Overfitting describes sensitivity to the training set. It does not mean that a model performs differently across groups.&lt;/p&gt;

&lt;h2&gt;
  
  
  Variance appears as a gap between training and unseen performance
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Variance&lt;/strong&gt; is a model’s sensitivity to the particular data on which it was trained. High variance appears as &lt;strong&gt;overfitting&lt;/strong&gt;: the model performs extremely well on its training data but poorly on unseen data.&lt;/p&gt;

&lt;p&gt;The model has learned details specific to the training set instead of learning a pattern that generalizes. The decisive evidence is the gap between training and unseen-data performance.&lt;/p&gt;

&lt;p&gt;Underfitting has a different signature:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Training performance&lt;/th&gt;
&lt;th&gt;Unseen-data performance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Underfitting, or high statistical bias&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overfitting, or high variance&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Healthy aggregate fit&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For example, suppose a model answers nearly every training example correctly but makes frequent errors on new examples. The strong training result does not show that the model is healthy. Combined with poor unseen performance, it shows overfitting.&lt;/p&gt;

&lt;p&gt;Now change the example: the model performs well on both training and unseen data, but its unseen-data accuracy is much lower for one demographic group. Neither underfitting nor overfitting fully describes that result. The model can be well fitted in aggregate while remaining unfair.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffg3hbx3mh38tf7gv8c8x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffg3hbx3mh38tf7gv8c8x.png" alt="Poor results on both sets mean underfitting; a large gap means overfitting; good results still require group checks." width="800" height="1336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Training and unseen performance separate underfitting from overfitting before fairness is assessed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A useful exam rule follows: when a question gives two performance figures—one for training data and one for unseen data—it is usually testing fitting. When it gives an overall figure and mentions a demographic group, it is testing fairness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dataset characteristics explain where disparities begin
&lt;/h2&gt;

&lt;p&gt;The exam separates four dataset characteristics because each answers a different question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inclusivity&lt;/strong&gt; asks whether the groups the system will serve are present at all. If a group has no records, the model cannot learn from it, and there are no rows on which to calculate that group’s performance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Diversity&lt;/strong&gt; asks whether the data represents the real range of cases, conditions, and contexts. A dataset may include every named group yet cover only typical cases, leaving the model unreliable at the edges.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Balance&lt;/strong&gt; asks whether represented groups appear in workable proportions. A group can be present but so heavily outnumbered that it has little effect on the training objective.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Curated sources&lt;/strong&gt; have known, selected, and documented origins. Curation makes the data’s provenance inspectable; it does not guarantee neutrality.&lt;/p&gt;

&lt;p&gt;Imagine a dataset containing 40,000 records from one group and 300 from another. The smaller group is present, so this is not an inclusivity failure. It is a balance failure: the group is represented but swamped.&lt;/p&gt;

&lt;p&gt;By contrast, if the second group has no records, “rebalance the dataset” is not yet a sufficient answer. There are no records to reweight. Data must first be collected from the missing group.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwc08syhhag5qsw2uknbm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwc08syhhag5qsw2uknbm.png" alt="Inclusivity checks presence, diversity checks range, balance checks proportion, and curation checks provenance." width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Dataset checks move from whether groups exist to whether their representation can support reliable learning.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Curation is especially easy to misread. A carefully documented dataset collected entirely from one narrow population can be curated and skewed at the same time. Curation helps you identify what the dataset represents; it does not make the representation fair.&lt;/p&gt;

&lt;p&gt;Bias may also enter through collection, labelling, training, deployment, or feedback. For example, inconsistent reviewers may assign different labels to similar cases. Rebalancing the dataset would not repair those prejudiced or inconsistent targets. The labels themselves require analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aggregate accuracy cannot establish fairness
&lt;/h2&gt;

&lt;p&gt;An aggregate metric is an average. It can describe total performance without revealing how that performance is distributed.&lt;/p&gt;

&lt;p&gt;Suppose a model reports 94% accuracy overall. That figure may be arithmetically correct while concealing much lower accuracy for a minority group. Nothing about the headline number answers the fairness question “for whom does the model work?”&lt;/p&gt;

&lt;p&gt;The appropriate technique is &lt;strong&gt;subgroup analysis&lt;/strong&gt;: calculate the same metric separately for each relevant group. Subgroup analysis is not an AWS product. It is the measurement habit that makes disparities visible.&lt;/p&gt;

&lt;p&gt;This produces an important sequence for exam scenarios:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Refuse to treat overall accuracy as fairness evidence.&lt;/li&gt;
&lt;li&gt;Identify the groups that need comparison.&lt;/li&gt;
&lt;li&gt;Recalculate the relevant metric for each group.&lt;/li&gt;
&lt;li&gt;Trace the disparity back through collection, labelling, training, deployment, and feedback.&lt;/li&gt;
&lt;li&gt;Select a detection or monitoring tool based on timing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Removing a demographic attribute is not a reliable fix. Other features may act as &lt;strong&gt;proxy variables&lt;/strong&gt;, meaning they carry information correlated with the removed attribute. Removing the attribute can also eliminate the field needed to calculate per-group metrics, making the disparity harder to detect without removing its cause.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fodhzz6o3n3mux13hcgpu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fodhzz6o3n3mux13hcgpu.png" alt="A healthy overall metric must be split by group before a disparity can be detected and traced." width="800" height="164"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Subgroup analysis exposes disparities that remain invisible inside a healthy overall result.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Human judgment still matters because somebody must decide which groups to compare. A metric can only expose disparities it was configured to measure. &lt;strong&gt;Human audits&lt;/strong&gt; review model behavior directly and can identify concerns that were never encoded in an automated check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Timing selects Clarify, Model Monitor, or A2I
&lt;/h2&gt;

&lt;p&gt;The AWS tools in these questions are not interchangeable. The simplest way to choose among them is to ask when the examination occurs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Amazon SageMaker Clarify&lt;/strong&gt; performs point-in-time bias measurement. It can assess bias in data before training and in a trained model after training. It also produces feature attributions, but in a bias question its relevant role is calculating bias metrics.&lt;/p&gt;

&lt;p&gt;Choose Clarify when the scenario asks whether a training dataset or trained model is biased now, particularly before deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SageMaker Model Monitor&lt;/strong&gt; watches a deployed model continuously for changes in data, model quality, and bias relative to a baseline.&lt;/p&gt;

&lt;p&gt;Choose Model Monitor when the scenario says the model was acceptable at launch but may have changed after deployment. Phrases such as “since launch,” “drift,” and “continually monitor” distinguish it from Clarify.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Amazon Augmented AI (Amazon A2I)&lt;/strong&gt; routes individual predictions to human reviewers at inference time, often when confidence is low.&lt;/p&gt;

&lt;p&gt;Choose A2I when a particular prediction or application needs human judgment. Do not choose it to calculate population-level disparity: A2I produces a human decision on a case, not a bias metric across groups.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fprtucwnsvektzs89jhh8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fprtucwnsvektzs89jhh8.png" alt="Clarify measures bias at a point in time, Model Monitor watches deployment, and A2I reviews individual cases." width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The point in time determines whether the answer is Clarify, Model Monitor, or A2I.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two other methods complete the exam’s detection set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analyzing label quality&lt;/strong&gt; checks whether labels are correct and consistently applied. Choose it when the evidence points to reviewers using inconsistent standards or when bias may be encoded in the training target.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human audits&lt;/strong&gt; are appropriate when the concern requires judgment that no configured metric captures. They are not an inferior substitute for automation; they examine blind spots created by the choice of what to measure.&lt;/p&gt;

&lt;p&gt;A compact selection rule is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario signal&lt;/th&gt;
&lt;th&gt;Best match&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Assess training data or a trained model now&lt;/td&gt;
&lt;td&gt;SageMaker Clarify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Watch for change after deployment&lt;/td&gt;
&lt;td&gt;SageMaker Model Monitor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Send one prediction for human review&lt;/td&gt;
&lt;td&gt;Amazon A2I&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compare performance across groups&lt;/td&gt;
&lt;td&gt;Subgroup analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Check inconsistent or prejudiced targets&lt;/td&gt;
&lt;td&gt;Label quality analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Investigate concerns no metric encoded&lt;/td&gt;
&lt;td&gt;Human audit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Common distractors fail because they answer the wrong question
&lt;/h2&gt;

&lt;p&gt;Several plausible answers recur in bias and variance scenarios.&lt;/p&gt;

&lt;p&gt;“Collect more data” does not fix missing coverage if the new records come from the same sources. More of the same population repeats the same skew. The corrective action must add relevant cases or groups.&lt;/p&gt;

&lt;p&gt;“Remove the demographic column” does not prevent discrimination when proxy variables remain. It may instead destroy the ability to measure the disparity.&lt;/p&gt;

&lt;p&gt;“Use Clarify for ongoing monitoring” ignores timing. Clarify measures at a point in time; Model Monitor handles continuous checks after deployment.&lt;/p&gt;

&lt;p&gt;“Use Model Monitor on the training dataset” places a deployment tool before deployment. There is no endpoint behavior to monitor yet.&lt;/p&gt;

&lt;p&gt;“Use A2I to measure bias” confuses case review with population analysis. A2I routes individual predictions to people; subgroup analysis or bias metrics reveal group-level disparity.&lt;/p&gt;

&lt;p&gt;“Call the group failure overfitting” confuses demographic performance with generalization. Overfitting requires strong training performance and poor unseen-data performance. A group disparity requires per-group evidence.&lt;/p&gt;

&lt;p&gt;Think of these distractors as category errors. Each proposed action may be useful somewhere, but it does not match the evidence, the level of analysis, or the point in the model lifecycle described by the question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Statistical bias means underfitting: the model performs poorly on training and unseen data.&lt;/li&gt;
&lt;li&gt;High variance means overfitting: training performance is strong, but unseen performance is poor.&lt;/li&gt;
&lt;li&gt;Societal bias means outcomes differ systematically across groups and can coexist with good aggregate performance.&lt;/li&gt;
&lt;li&gt;Inclusivity is about presence; balance is about proportion; diversity is about range; curation is about documented provenance.&lt;/li&gt;
&lt;li&gt;Overall accuracy cannot establish fairness. Calculate metrics separately for relevant groups.&lt;/li&gt;
&lt;li&gt;SageMaker Clarify measures data or model bias at a point in time.&lt;/li&gt;
&lt;li&gt;SageMaker Model Monitor watches a deployed model for drift in quality and bias.&lt;/li&gt;
&lt;li&gt;Amazon A2I routes individual predictions to human reviewers; it is not a population bias metric.&lt;/li&gt;
&lt;li&gt;Label quality analysis examines biased or inconsistent targets, while human audits catch concerns no metric was configured to detect.&lt;/li&gt;
&lt;li&gt;Timing, not the general word “bias,” is the strongest clue when selecting an AWS tool.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These rules settle how to distinguish the concepts and select the exam’s named detection methods. They do not prove that a real model is fair: that still depends on which groups were examined, whether the labels and data reflect the deployment population, and whether human reviewers looked for harms the chosen metrics could not express.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>certification</category>
    </item>
    <item>
      <title>Your LLM Prompt Costs More Than You Think — Build a Token &amp; Cost Calculator in Colab</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:59:45 +0000</pubDate>
      <link>https://dev.to/confident_prep/your-llm-prompt-costs-more-than-you-think-build-a-token-cost-calculator-in-colab-23d8</link>
      <guid>https://dev.to/confident_prep/your-llm-prompt-costs-more-than-you-think-build-a-token-cost-calculator-in-colab-23d8</guid>
      <description>&lt;p&gt;You don't need an API key to find out your AI feature is quietly expensive.&lt;/p&gt;

&lt;p&gt;You don't need a production account.&lt;/p&gt;

&lt;p&gt;You don't need to wait for the first bill.&lt;/p&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google Colab&lt;/li&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pip install tiktoken&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;About 10 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last article in this series, "Give Your Chatbot a Memory in Google Colab Before Your Next AI Interview," was about managing a conversation once it gets long. &lt;em&gt;(Link it to that post's live Dev.to URL when you paste this in — it isn't in this repo.)&lt;/em&gt; This one rewinds further — before memory, before RAG, before agents — to the mechanism everything else in an LLM system is built on top of: tokens, and what they actually cost you. If you can't answer "how many tokens is this, and what does that mean in dollars," the rest of the stack doesn't matter yet.&lt;/p&gt;

&lt;p&gt;Let's measure it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Are We Building?
&lt;/h2&gt;

&lt;p&gt;Every LLM call is billed and bounded by tokens, not words or characters. We'll build three small tools:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A real tokenizer that counts tokens the way a production model actually would&lt;/li&gt;
&lt;li&gt;A cost calculator that turns a token count into a daily and yearly dollar figure&lt;/li&gt;
&lt;li&gt;A tiny simulation of what "temperature" actually does to next-token sampling&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No fine print, no mocked numbers — every count below came from running the code, not estimating it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Count Tokens for Real
&lt;/h2&gt;

&lt;p&gt;Skip guessing "about 4 characters per token." Use the same tokenizer family production APIs use.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tiktoken&lt;/span&gt;

&lt;span class="n"&gt;enc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tiktoken&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_encoding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cl100k_base&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;text_a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The meeting is scheduled for Thursday at 3pm. Please bring your notes from the last session and any open questions you want to discuss with the team.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;text_b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The rate-limiting factor in transformer inference is memory bandwidth during KV-cache reads, not FLOPs. Prefill is compute-bound; autoregressive decode is memory-bandwidth-bound.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;text_c&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Here is the function signature: async def fetch_user(user_id: UUID, db: AsyncSession) -&amp;gt; Optional[UserModel]:. It should return None if the user does not exist, not raise an exception.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A - plain English&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text_a&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;B - technical jargon&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text_b&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;C - mixed code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text_c&lt;/span&gt;&lt;span class="p"&gt;)]:&lt;/span&gt;
    &lt;span class="n"&gt;toks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;enc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;toks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; words -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;toks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens/word&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output, run just now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A - plain English: 31 tokens, 27 words -&amp;gt; 1.15 tokens/word
B - technical jargon: 37 tokens, 21 words -&amp;gt; 1.76 tokens/word
C - mixed code: 43 tokens, 27 words -&amp;gt; 1.59 tokens/word
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same rough word count across all three (21–27 words), but Text B costs &lt;strong&gt;53% more tokens per word&lt;/strong&gt; than Text A. Compound terms like "rate-limiting" get chopped into multiple sub-word pieces; UUIDs and type annotations in Text C do the same. Plain English is the cheapest register you can write in — technical prompts are a variable tax, not a fixed multiplier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Turn Tokens Into Dollars
&lt;/h2&gt;

&lt;p&gt;A token count is abstract until it's attached to a query volume and a price sheet.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;PRICING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;3.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;15.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;   &lt;span class="c1"&gt;# $ per 1M tokens
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-haiku&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;4.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;estimate_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;queries_per_day&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;days&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;365&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;enc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;daily_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;queries_per_day&lt;/span&gt;
    &lt;span class="n"&gt;daily_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;daily_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PRICING&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_cost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_cost&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;days&lt;/span&gt;

&lt;span class="n"&gt;system_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a support assistant for Acme Cloud. Always answer in a friendly, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;professional tone. Never reveal internal pricing formulas. If the user asks &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;about billing, refer them to the billing dashboard at app.acme.com/billing. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;If the user reports an outage, check the status page before responding, and &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;always include the current incident ID if one is open. Keep responses under &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;150 words unless the user explicitly asks for more detail.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-haiku&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_tok&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_cost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;yearly_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimate_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens/call, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;daily_tok&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens/day, $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;daily_cost&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/day, $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;yearly_cost&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/year&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Real output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;claude-sonnet: 87 tokens/call, 87,000 tokens/day, $0.26/day, $95/year
claude-haiku: 87 tokens/call, 87,000 tokens/day, $0.07/day, $25/year
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One system prompt, sent unchanged on every one of 1,000 daily queries. $95/year on Sonnet doesn't sound alarming — until this is one of six prompts in your app, or volume is 100,000/day instead of 1,000. The formula doesn't change; the multiplier does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Cut the Bill Without Cutting the Feature
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;trimmed_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are Acme Cloud&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s support assistant. Friendly, professional tone. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Never reveal pricing formulas. Billing questions -&amp;gt; app.acme.com/billing. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Outage reports -&amp;gt; check status page, include open incident ID. Max 150 words.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_cost2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;yearly_cost2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimate_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;trimmed_prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Trimmed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens/call (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="mi"&gt;87&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;% shorter), $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;daily_cost2&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/day, $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;yearly_cost2&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;,.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/year&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Trimmed: 47 tokens/call (46% shorter), $0.14/day, $51/year
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Denser prose cut the prompt nearly in half and the yearly cost from $95 to $51 — same rules enforced, zero functionality lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: What Temperature Actually Does
&lt;/h2&gt;

&lt;p&gt;Temperature doesn't make a model smarter or dumber — it reshapes how sharply it commits to its top guess. Simulate it with a toy 5-word distribution over "The weather today is ___":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;
&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;candidates&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sunny&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cloudy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cold&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unpredictable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;magnificent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;logits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;     &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;2.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="mf"&gt;2.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;softmax_with_temperature&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1e-6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;logits&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scaled&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;exps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;scaled&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exps&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;exps&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;cum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;cum&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;cum&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;temp&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;softmax_with_temperature&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;temp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;picks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temp=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;temp&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: probs=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  8 samples: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;picks&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;temp=0.2: probs={'sunny': 0.71, 'cloudy': 0.26, 'cold': 0.02, 'unpredictable': 0.0, 'magnificent': 0.0}
  8 samples: ['sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny', 'sunny']

temp=1.0: probs={'sunny': 0.36, 'cloudy': 0.3, 'cold': 0.18, 'unpredictable': 0.11, 'magnificent': 0.05}
  8 samples: ['sunny', 'cloudy', 'sunny', 'sunny', 'cloudy', 'cold', 'sunny', 'sunny']
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same logits both times. At &lt;code&gt;temp=0.2&lt;/code&gt; the distribution collapses onto "sunny" — 8 for 8. At &lt;code&gt;temp=1.0&lt;/code&gt; the same model "opinion" produces four different words across 8 draws. Nothing about the model's knowledge changed — only how willing it is to pick outside its top guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make It Reusable
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;tokenize_and_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;queries_per_day&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;daily_cost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;yearly_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;estimate_cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;queries_per_day&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tokens_per_call&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily_cost_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;daily_cost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;yearly_cost_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;yearly_cost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One function, three inputs (text, volume, model), a real dollar figure out. Swap in your own system prompt before your next design review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Colab Experiments to Try
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Paste your own longest system prompt into Step 1 — is its tokens/word ratio closer to plain English or technical jargon?&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;estimate_cost&lt;/code&gt; at 10,000 queries/day. Watch the yearly figure move from "line item" to "budget conversation."&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;gpt-4o-mini&lt;/code&gt; (&lt;code&gt;$0.15&lt;/code&gt; input) to &lt;code&gt;PRICING&lt;/code&gt; and compare its yearly cost against &lt;code&gt;claude-haiku&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Change &lt;code&gt;logits&lt;/code&gt; so all 5 candidates are nearly equal, then re-run both temperatures — does low temperature still look deterministic?&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Interview Questions Hidden Inside This Notebook
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does a technical or code-heavy prompt cost more than plain English at the same word count?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sub-word tokenization splits compound technical terms and punctuation-heavy syntax (UUIDs, type hints) into more pieces than common English words. Word count and token count are only loosely correlated — measure tokens directly, never estimate from word count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What lever actually reduces LLM cost on a fixed feature set?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Prompt density. The Step 3 rewrite cut cost 46% by removing redundant phrasing, not functionality. Token count is a writing-quality problem before it's an infrastructure problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does temperature affect what the model "knows"?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. It reshapes sampling over an unchanged probability distribution — it doesn't touch the model's underlying weights. Low temperature for tasks needing the same answer every time; higher temperature only where variety is the actual goal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Calculator Is the Easy Part
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;tokenize_and_cost()&lt;/code&gt; is maybe six lines. A real production cost model also has to account for output tokens (billed 3-5x the input rate, harder to bound), prompt caching (steep discounts on repeated system prompts this notebook doesn't model), and retries or tool-call round-trips that resend and re-bill context.&lt;/p&gt;

&lt;p&gt;Interviewers ask about token cost because a candidate who can name a real number, on the spot, has clearly built something — not just read about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to Go Deeper?
&lt;/h2&gt;

&lt;p&gt;This notebook covers the mechanics. The full session covers the "how does an LLM generate a response?" 5-beat answer end to end, including the hallucination root cause this notebook didn't touch:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/" rel="noopener noreferrer"&gt;https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There's also a companion piece, &lt;em&gt;"How Does an LLM Actually Generate a Response?" — The Interview Question Everyone Gets 80% Right&lt;/em&gt;, with 8 interview questions pulled straight from this same session. &lt;em&gt;(Link it to that post's Dev.to URL once both are published.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Measure your own prompt before your next standup. A real number beats a guess — in a cost review and in an interview.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>"How Does an LLM Actually Generate a Response?" — The Interview Question Everyone Gets 80% Right</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:59:23 +0000</pubDate>
      <link>https://dev.to/confident_prep/how-does-an-llm-actually-generate-a-response-the-interview-question-everyone-gets-80-right-3eed</link>
      <guid>https://dev.to/confident_prep/how-does-an-llm-actually-generate-a-response-the-interview-question-everyone-gets-80-right-3eed</guid>
      <description>&lt;p&gt;If you're prepping for an AI engineering interview, this question is coming: &lt;strong&gt;"How does an LLM generate a response?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most candidates get the first two beats right — tokens, embeddings — then start hand-waving. The ones who land the offer name all five beats without prompting, including the one most people skip.&lt;/p&gt;

&lt;p&gt;There's a companion piece, &lt;em&gt;"Your LLM Prompt Costs More Than You Think"&lt;/em&gt;, that builds a token/cost calculator in Colab you can run in 10 minutes with zero API key. &lt;em&gt;(Link it here once both are published.)&lt;/em&gt; This one is the pure Q&amp;amp;A version — no code, just the questions and the answers that separate a senior response from a junior one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core question, and the 5 beats interviewers listen for
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"How does an LLM generate a response?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A strong 90-second answer hits all five:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tokens&lt;/strong&gt; — sub-word units; the model reads these, not words&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings&lt;/strong&gt; — each token becomes a vector of numbers capturing meaning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attention&lt;/strong&gt; — scores relevance across every token in context, for each new token generated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation / sampling&lt;/strong&gt; — next-token prediction, one token at a time, until done&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hallucination root cause&lt;/strong&gt; — plausibility ≠ truth; there's no truth-check step in the loop&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Miss beat 5 and expect a direct follow-up. It's the most commonly skipped beat, and interviewers know it.&lt;/p&gt;

&lt;h2&gt;
  
  
  8 questions from this session, answered
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. What is a token, and why does it matter for your job?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A sub-word chunk — the model's basic reading unit. Not a word, not a character. "unbelievable" is 3 tokens. "ChatGPT" is 3 tokens. It matters because you pay per token, not per word, and the model has a hard limit on how many it can process per request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. What's the difference between "context" and the "context window"?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Context is everything the model can see in one conversation — your messages, documents, its own replies. The context window is the size limit on that everything: the maximum token count (input plus output) per request. It resets completely on every new call — not persistent memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. What is an embedding, in plain terms?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A list of numbers representing a token's meaning and usage, not its spelling. Similar words land at nearby coordinates; unrelated words land far apart. The same word can carry different embeddings depending on the sentence it's in ("bank" the institution vs. "bank" the river).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. What does attention actually do?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For every token about to be generated, attention scores every other token in context for relevance. High-relevance tokens get more influence over what comes next. This is also why a cluttered, noisy prompt produces worse output — irrelevant tokens compete for attention against the instruction that actually matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Why does an LLM hallucinate — what's the real root cause?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model predicts statistically plausible text. There is no truth-check step anywhere in generation. Confident, authoritative writing was abundant in training data, so the model learned confident tone as a default style — a style that shows up whether or not the underlying claim is true. Plausibility is not truth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. What does temperature control, and when do you use low vs. high?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;How peaked or flat the next-token probability distribution is. Low (0.0–0.3): the model almost always picks its top-probability token — use for fact extraction, structured output, anything that needs to be the same every call. High (0.7–1.0): more willing to pick lower-probability tokens — use for brainstorming or creative variation. Temperature never changes what the model knows, only how consistently it says it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. How does a chatbot "remember" earlier messages if the model resets every call?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It doesn't, on its own. The application layer replays prior turns back into the context window on every new request. The continuity a user experiences is maintained by the app, not the model. Strip out that replay logic and the model has never heard of the user before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;8. Why can't an LLM answer questions about your company's internal documents?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Training data is frozen at a cutoff date, and internal documents were never in it to begin with. Anything private or created after the cutoff has to be explicitly supplied in context — which is the entire premise behind retrieval-augmented generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What separates a strong answer from a weak one
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Say:&lt;/strong&gt; "Tokens, not words." "Embeddings capture meaning as vectors." "Attention scores relevance across the full context." "Plausibility, not truth — there's no truth-check step." "Context window is a fixed buffer that resets."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Avoid:&lt;/strong&gt; "It's just autocomplete" (technically true, reads as surface-level). "It thinks like a human" (signals a real misconception). "It searches the internet" (only true with retrieval tools wired in). "It hallucinates because it lacks data" (names the wrong root cause — it hallucinates even &lt;em&gt;with&lt;/em&gt; the data, because nothing checks the output against it).&lt;/p&gt;

&lt;h2&gt;
  
  
  Want the full 5-beat model answer, word for word?
&lt;/h2&gt;

&lt;p&gt;The course session covers this question with a complete scored model answer, the do's/don'ts list interviewers are actually listening for, and the hands-on Colab notebook that makes tokens and cost tangible before you ever sit down for the interview:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/" rel="noopener noreferrer"&gt;https://confidentprep.com/courses/ai-ml-for-interview/1-how-llms-actually-work/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Say your answer out loud before your next interview. Reading it silently and being able to say it under pressure are two different skills.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>interview</category>
      <category>career</category>
    </item>
    <item>
      <title>Walk Me Through How You'd Design a Production Prompt — 9 Interview Questions, Answered</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:57:35 +0000</pubDate>
      <link>https://dev.to/confident_prep/walk-me-through-how-youd-design-a-production-prompt-9-interview-questions-answered-2k9a</link>
      <guid>https://dev.to/confident_prep/walk-me-through-how-youd-design-a-production-prompt-9-interview-questions-answered-2k9a</guid>
      <description>&lt;p&gt;There's one question that opens almost every AI engineering interview's prompt-design section:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Walk me through how you'd design a production prompt for a structured extraction task."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It sounds open-ended. It isn't. Interviewers are listening for five specific beats, and the follow-up questions below are how they check for the ones you skip. This is the companion piece to &lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/2-prompt-engineering-and-the-llm-api-surface/" rel="noopener noreferrer"&gt;the hands-on Colab notebook&lt;/a&gt; where we measured an 80%-to-0% format-drift swing from adding few-shot examples — if you want the receipts behind question 3 below, that's where they are.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What's the difference between a system prompt and a user message?
&lt;/h2&gt;

&lt;p&gt;The system prompt is developer-controlled — role, rules, output format — and runs invisibly on every call. The user message is the runtime task or input, changing with each call. Miss this distinction and you'll bake per-request data into a prompt that should be static, or vice versa.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What does RCTF stand for, and what happens if you skip one letter?
&lt;/h2&gt;

&lt;p&gt;Role, Context, Task, Format. Skip Format and the model chooses its own output shape — usually inconsistently. Skip Context and it fabricates facts it doesn't have (product names, plan tiers). Skip Role and tone drifts call to call. Skip Task and you get a summary when you wanted a classification. Interviewers who ask "what happens if you skip X" are testing whether you understand &lt;em&gt;why&lt;/em&gt; each field exists, not just that it exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What temperature would you use for a support ticket classifier, and why?
&lt;/h2&gt;

&lt;p&gt;Zero, or close to it. Classification feeds a parser, and parsers need identical output for identical input. Temperature above zero introduces sampling randomness — and that randomness doesn't just reword the answer, it can surface as a genuinely different output shape. Say the number if you have it: a naive zero-shot classifier can drift to a different JSON schema over 80% of the time at moderate temperature. Anchoring the same call with a few worked examples took that to zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. When do you reach for few-shot instead of zero-shot?
&lt;/h2&gt;

&lt;p&gt;Always start zero-shot — it's cheaper and easier to maintain. Move to few-shot only when zero-shot's &lt;em&gt;content&lt;/em&gt; is right but the &lt;em&gt;format&lt;/em&gt; is wrong or inconsistent. Few-shot examples teach output shape, not new facts. If the content itself is wrong, more examples won't fix it — that's a Context problem, not a technique problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. What's the difference between few-shot prompting and fine-tuning?
&lt;/h2&gt;

&lt;p&gt;Few-shot: examples live in the prompt, inference-time only, weights unchanged, cheap to update. Fine-tuning: the model is retrained on your dataset, weights change permanently, requires thousands of examples and real infrastructure. Try few-shot first. Fine-tuning is for stable, high-volume, narrowly-defined tasks where prompting has genuinely been exhausted — not a first instinct.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. When does chain-of-thought help, and what does it cost you?
&lt;/h2&gt;

&lt;p&gt;CoT helps on multi-condition or multi-step tasks where jumping straight to an answer causes errors — escalation decisions, trade-off analysis. The instruction "think step by step" forces intermediate reasoning tokens that improve the final answer. The cost: more tokens, more latency. Don't reach for it on simple single-condition tasks — that's needless spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. What is prompt injection, and how do you actually prevent it?
&lt;/h2&gt;

&lt;p&gt;User-supplied text overwrites developer instructions because, to the model, both are just tokens — there's no privileged execution layer separating "your rules" from "their input." The fix that holds up: XML-delimit user input (&lt;code&gt;&amp;lt;user_input&amp;gt;...&amp;lt;/user_input&amp;gt;&lt;/code&gt;) and instruct the model to treat that block as data only. The fix that doesn't hold up: asking nicely in the system prompt and hoping. A system prompt is a request, not a boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Why does a model confidently answer with stale stock prices or outdated facts?
&lt;/h2&gt;

&lt;p&gt;Knowledge cutoff. The model has no internet access and no live retrieval — it generates the statistically most plausible continuation from training data, which may be months or years stale, with the same confident tone as a correct answer. There's no "I don't know current data" reflex unless you build one. The fix is injecting current data into context directly (RAG), never trusting model memory for anything time-sensitive.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. What does max_tokens actually control, and what breaks if you set it wrong?
&lt;/h2&gt;

&lt;p&gt;The length of the &lt;em&gt;response&lt;/em&gt;, not the input — that's the context window's job. Set too low, structured output truncates mid-JSON and your parser throws. Set too high, you're paying for and waiting on output nobody asked for. Rule of thumb: roughly 2x your expected output length, tighter (80-150) for constrained JSON classifiers.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;The beat interviewers probe hardest:&lt;/strong&gt; failure modes. A candidate who nails RCTF and parameters but never mentions injection, knowledge cutoff, or format validation unprompted gets asked "what could go wrong?" — and that follow-up is where senior and junior answers actually separate.&lt;/p&gt;

&lt;p&gt;Want to see the 80%-to-0% drift number measured live, not just quoted? &lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/2-prompt-engineering-and-the-llm-api-surface/" rel="noopener noreferrer"&gt;The hands-on notebook is here&lt;/a&gt; — zero API key required, about 12 minutes.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>interview</category>
      <category>career</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Are You Ready for Your Job? Try to Answer These Questions.</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:55:44 +0000</pubDate>
      <link>https://dev.to/confident_prep/are-you-ready-for-your-job-try-to-answer-these-questions-2p79</link>
      <guid>https://dev.to/confident_prep/are-you-ready-for-your-job-try-to-answer-these-questions-2p79</guid>
      <description>&lt;p&gt;Not a quiz for fun. A gut check.&lt;/p&gt;

&lt;p&gt;These are three questions that show up in real AI engineering interviews — not "define RAG" trivia, but the scenario-style questions interviewers actually use to tell a candidate who's built things apart from one who's only read about them. Read each one, answer it out loud before you read the framework underneath, and be honest about how close you got.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🎯 &lt;strong&gt;TL;DR:&lt;/strong&gt; Three real interview questions, no trivia. Answer each one out loud before reading the framework. If you reach for a definition instead of a process, that's the exact gap other candidates are closing right now — not by reading more, but by building.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  🩹 1. "Your RAG chatbot just gave a customer a confident, completely wrong answer. Walk me through how you'd find out why — and stop it from happening again."
&lt;/h2&gt;

&lt;p&gt;This is a debugging question disguised as a RAG question. Interviewers ask it because "explain how RAG works" tells them you read an article, but this tells them whether you've actually operated one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to answer it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Split the failure in two before you diagnose anything: was it a &lt;strong&gt;retrieval&lt;/strong&gt; problem (the right chunk never made it into context) or a &lt;strong&gt;generation&lt;/strong&gt; problem (the right chunk was there and the model ignored it anyway)? Say this split out loud first — it's the single biggest signal you know what you're doing.&lt;/li&gt;
&lt;li&gt;Describe how you'd check retrieval: log the retrieved chunks alongside the query, and eyeball whether the answer's source text was even in there.&lt;/li&gt;
&lt;li&gt;Describe how you'd check generation: if the source was there but the model still got it wrong, that's a prompt/grounding problem — talk about instructing the model to cite or refuse when the context doesn't support an answer.&lt;/li&gt;
&lt;li&gt;Close with prevention, not just the one-time fix: a small eval set of known-answer questions you can re-run after every prompt or retrieval change, so this doesn't quietly regress next sprint.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  ⚖️ 2. "We could fix this with a better prompt, or by fine-tuning the model. How do you decide which one?"
&lt;/h2&gt;

&lt;p&gt;This question is a trap for candidates who only know one hammer. It's testing judgment, not a definition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to answer it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lead with prompt engineering as the default: it's fast, cheap, needs no training data, and you can iterate on it in minutes.&lt;/li&gt;
&lt;li&gt;Name the specific conditions that push you toward fine-tuning instead: you have a large, stable, labeled dataset; you need consistent output formatting or tone at a scale prompting can't reliably hold; or your prompt is already maxed out on context budget and instructions are still being ignored.&lt;/li&gt;
&lt;li&gt;Mention the cost you're trading away: fine-tuning is slower to iterate and harder to reverse than a prompt change. Say that out loud — it shows you're weighing the decision, not just picking a technique.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🚨 3. "The feature demos perfectly, but real users are breaking it in ways you never saw in testing. What's your process?"
&lt;/h2&gt;

&lt;p&gt;This is the question that separates "I built a demo" from "I've operated something in production."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to answer it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Start with the gap between demo inputs and real inputs: real users send empty strings, five-paragraph pastes, other languages, and adversarial prompts your test set never covered.&lt;/li&gt;
&lt;li&gt;Talk about what you need logged &lt;em&gt;before&lt;/em&gt; this happens, not after: full request/response pairs, retrieved context, any tool calls — without that, you're debugging blind.&lt;/li&gt;
&lt;li&gt;Mention a containment step: a feature flag or circuit breaker so you can turn the feature off or roll back the version for affected users while you dig in, instead of leaving it broken in front of customers while you investigate.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🪞 How did you do?
&lt;/h2&gt;

&lt;p&gt;If you answered all three with that level of specificity, unprompted — that's rare, and it means you've actually operated systems like this before.&lt;/p&gt;

&lt;p&gt;If you found yourself reaching for a textbook definition instead of a process, that's useful information. It means you understand the concepts but haven't yet built the judgment that only comes from shipping and breaking a few of these yourself. &lt;strong&gt;Other candidates in your interview pool are closing that exact gap right now — not by reading more, but by building.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That judgment isn't something an article can hand you — including this one. It comes from building the actual systems these questions are about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: knowing the definition doesn't get the offer. Having broken the thing once does.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://confidentprep.com/paths" rel="noopener noreferrer"&gt;See the paths that build that judgment →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>interview</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>career</category>
    </item>
    <item>
      <title>Tokens, Embeddings and the Foundation Model Lifecycle for AWS AIF-C01</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:54:57 +0000</pubDate>
      <link>https://dev.to/confident_prep/tokens-embeddings-and-the-foundation-model-lifecycle-for-aws-aif-c01-3bo9</link>
      <guid>https://dev.to/confident_prep/tokens-embeddings-and-the-foundation-model-lifecycle-for-aws-aif-c01-3bo9</guid>
      <description>&lt;p&gt;An exam question describes a long document, a search request and a model being adapted for a specialist task. Which detail matters first: the words, the vectors or the training stage? For AWS AIF-C01, the reliable answer comes from separating three layers: tokens are what a foundation model processes, embeddings make meaning comparable, and the foundation model lifecycle turns a broadly capable model into a callable system that can improve through feedback. Once those layers are clear, many plausible distractors stop looking plausible.&lt;/p&gt;

&lt;h2&gt;
  
  
  A foundation model is a base, not a finished solution
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;foundation model&lt;/strong&gt; is a large model pre-trained on broad, general data that can be adapted to many downstream tasks without being trained specifically for each one. Its defining property is adaptability, not size.&lt;/p&gt;

&lt;p&gt;Think of it as a general-purpose kitchen. The kitchen supports many dishes, but it is not itself a finished meal. Instructions, selected ingredients and sometimes specialist preparation are still needed for a particular result. In the same way, a foundation model supplies broad capability that can later be directed or adapted.&lt;/p&gt;

&lt;p&gt;That distinction matters because an exam question may describe a very large model and invite you to classify it as a foundation model on size alone. Size is insufficient. Look for broad pre-training plus the ability to support multiple downstream tasks.&lt;/p&gt;

&lt;p&gt;A transformer-based large language model is one possible foundation model. A &lt;strong&gt;transformer&lt;/strong&gt; relates every token in an input to every other token, allowing context to affect meaning. Consider the word “bank”: the surrounding tokens distinguish a river bank from a savings bank. The model predicts an output from those relationships; it does not behave like a database looking up a stored sentence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqw73r4eu8ht8k0uvyf0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faqw73r4eu8ht8k0uvyf0.png" alt="Broad data produces a foundation model that supports several forms of adaptation." width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Broad pre-training creates a base that can be adapted to different downstream tasks.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This gives you the first exam rule: define the thing before choosing its use. If the question asks what distinguishes a foundation model, answer with broad pre-training and adaptability. Do not substitute a use case such as summarization, and do not rely on “large” as the decisive clue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokens are the units the model actually receives
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;token&lt;/strong&gt; is a learned sub-word unit: a chunk of text that a model processes. It is neither necessarily a word nor necessarily a character.&lt;/p&gt;

&lt;p&gt;Common words may remain whole tokens. Less common words can split into several pieces, while identifiers containing digits and punctuation may fragment further. That is why a sentence’s word count does not tell you its token count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; compare an ordinary word such as &lt;code&gt;the&lt;/code&gt; with an identifier such as &lt;code&gt;AIF-C01&lt;/code&gt;. The common word may be represented as one token, while the identifier can become several because its letters, digits and punctuation may split. The exact split depends on the tokenizer, but the exam-relevant principle does not: one word does not reliably equal one token.&lt;/p&gt;

&lt;p&gt;Before the model processes text, a tokenizer converts it into token IDs, which are integers. The model operates on those IDs rather than directly reading letters as a person does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnz8b64mprc1353xi1z9j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnz8b64mprc1353xi1z9j.png" alt="Raw text becomes sub-word tokens, token IDs and finally model output tokens." width="800" height="492"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A tokenizer turns text into sub-word tokens and then into the integer IDs a model processes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Tokens also apply in both directions. The input consumes tokens, and the generated output consists of tokens too. In a continuing conversation, prior turns may be included again as input, so the amount processed can grow as the conversation grows.&lt;/p&gt;

&lt;p&gt;This exposes two common distractors. The first treats tokens as words and assumes a direct one-to-one conversion. The second counts only the prompt and forgets the output. Reject both by asking, “What did the model receive, and what did it emit?”&lt;/p&gt;

&lt;p&gt;A rough English-language guide is about four characters per token, but it is only a rule of thumb. Code, identifiers and many languages behave differently. Use the approximation only when a question clearly asks for rough reasoning; never promote it into an exact formula.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunking controls what can be processed and retrieved
&lt;/h2&gt;

&lt;p&gt;Long documents create two separate problems. A model can process only a bounded amount at once, and retrieval should return the relevant passage rather than an entire manual. &lt;strong&gt;Chunking&lt;/strong&gt; addresses both by splitting a document into smaller pieces before processing or storage.&lt;/p&gt;

&lt;p&gt;Think of it as dividing a book into chapters and building an index. When someone asks about one procedure, the useful response is the relevant chapter or passage—not the entire library.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; suppose a collection contains lengthy case documents, and a reader wants the passages related to a described situation. The documents should first be divided into chunks. Those chunks can then be compared with the reader’s query so that only the most relevant pieces are supplied for further processing.&lt;/p&gt;

&lt;p&gt;Chunk size introduces a tradeoff. A chunk that is too large may include the answer but surround it with irrelevant material. A chunk that is too small may retrieve a sentence whose necessary context was left in a neighbouring chunk.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpfnyrxlp722oypotyqw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpfnyrxlp722oypotyqw.png" alt="Large chunks add irrelevant material, while small chunks can lose necessary context." width="798" height="238"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Useful chunks balance enough context against the irrelevant material returned with an answer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Chunking is a preparation step, not a model family. If an exam scenario combines long documents, semantic search and summarization, do not force one technique to solve everything. Chunking divides the material; embeddings locate relevant chunks; a language model can summarize what was found.&lt;/p&gt;

&lt;h2&gt;
  
  
  Embeddings turn meaning into a comparable position
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;vector&lt;/strong&gt; is a fixed-length ordered list of numbers. By itself, that definition says nothing about meaning. An &lt;strong&gt;embedding&lt;/strong&gt; is a vector produced so that its position represents meaning. Every embedding is a vector, but not every vector is an embedding.&lt;/p&gt;

&lt;p&gt;Think of embeddings as map coordinates. Coordinates do not contain a written description of a city; they place it relative to other locations. Likewise, an embedding does not contain a readable summary of its source text. It places that text in a vector space where distance can represent relatedness.&lt;/p&gt;

&lt;p&gt;This is why embeddings cannot be decoded as though they were compressed documents. Their useful operation is comparison. When two embeddings occupy nearby positions, their source material is treated as meaningfully related.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; keyword search may fail to connect “cheap flights” with “budget airfare” because the phrases share no word. Embedding search can place them near each other because their meanings are related.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffi2gqsps68c2j55vr1zs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffi2gqsps68c2j55vr1zs.png" alt="Related phrases occupy nearby embedding positions while an unrelated phrase is farther away." width="800" height="330"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Embeddings place related meanings near one another even when their wording differs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That same comparison supports several capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Semantic search finds material with related meaning rather than matching only strings.&lt;/li&gt;
&lt;li&gt;Recommendation finds items near those associated with a preference.&lt;/li&gt;
&lt;li&gt;Clustering groups documents whose positions are close.&lt;/li&gt;
&lt;li&gt;Retrieval selects the chunks most related to a question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Search and recommendation are therefore comparison problems, not generation problems. A language model may generate a final response after retrieval, but generation is not what identifies the nearest material.&lt;/p&gt;

&lt;p&gt;On the exam, “vector representation of meaning” is a good start but not a complete explanation. Add that position encodes meaning, distance enables comparison, and the original text cannot simply be read back out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokens and embeddings do different jobs in one system
&lt;/h2&gt;

&lt;p&gt;Tokens and embeddings are easy to blur because both transform text into numbers. Their purposes, however, are different.&lt;/p&gt;

&lt;p&gt;Tokens are the units a model processes in sequence. Embeddings are positions used to compare meaning. Tokenization answers, “What units enter the model?” Embedding answers, “How can relatedness be measured?”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; consider a reader asking a natural-language question about a long policy document. The document is divided into chunks. Each chunk and the question receive embeddings. Their positions are compared to retrieve the closest chunk. That retrieved text is then tokenized when it is supplied to a language model, which produces an answer as output tokens.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm4tklfowodwcn2txlaqs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm4tklfowodwcn2txlaqs.png" alt="A document is chunked, compared through embeddings and passed as tokens for generation." width="800" height="106"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Chunking prepares text, embeddings retrieve by meaning, and tokens carry text through generation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A distractor may describe an embedding as a compact summary that the language model reads directly. That collapses two jobs into one. Keep the pipeline explicit: embeddings support comparison; retrieved text supplies content; tokenization converts that text into the units the model processes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lifecycle explains how a foundation model becomes usable
&lt;/h2&gt;

&lt;p&gt;The &lt;strong&gt;foundation model lifecycle&lt;/strong&gt; has seven stages in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Data selection produces a chosen, scoped corpus.&lt;/li&gt;
&lt;li&gt;Model selection produces a named base model.&lt;/li&gt;
&lt;li&gt;Pre-training produces a model with general capability.&lt;/li&gt;
&lt;li&gt;Fine-tuning produces a model adapted to a task or domain.&lt;/li&gt;
&lt;li&gt;Evaluation produces a pass or fail against a defined bar.&lt;/li&gt;
&lt;li&gt;Deployment produces a callable model.&lt;/li&gt;
&lt;li&gt;Feedback produces signals from real use.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The most important ordering relationship is that pre-training precedes fine-tuning. General capability must exist before that capability can be adapted to a narrower task or domain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; imagine preparing a model for specialist document work. The relevant corpus is selected, and a base model is chosen. Pre-training supplies broad capability; fine-tuning adapts the model. Evaluation checks it against an established bar, deployment makes it callable, and feedback reveals where another adjustment may be needed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06f59zqrdp2bbxnu1tt4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06f59zqrdp2bbxnu1tt4.png" alt="Seven lifecycle stages run in order, with feedback returning to fine-tuning." width="799" height="88"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The lifecycle builds general capability, adapts it, deploys it and loops feedback into improvement.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The feedback loop is not optional decoration. Without it, the diagram is merely a one-way release sequence. Feedback is a lifecycle stage because evidence from real use can drive another round of adaptation and evaluation.&lt;/p&gt;

&lt;p&gt;Also separate the model lifecycle from a broader project pipeline. A project may begin with a business goal and include operational monitoring. The foundation model lifecycle follows the model itself: its data, base selection, general training, adaptation, evaluation, deployment and feedback.&lt;/p&gt;

&lt;p&gt;For ordering questions, memorizing seven labels is not enough. Attach each stage to its output. That makes inversions easier to spot: a task-adapted model cannot be the output of pre-training, and a callable model cannot precede deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Exam signals reveal which concept is being tested
&lt;/h2&gt;

&lt;p&gt;When a scenario feels crowded, identify the verb and the object:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal in the question&lt;/th&gt;
&lt;th&gt;Concept to test&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Split a long document&lt;/td&gt;
&lt;td&gt;Chunking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processed unit or input/output count&lt;/td&gt;
&lt;td&gt;Tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Match meaning rather than exact words&lt;/td&gt;
&lt;td&gt;Embeddings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adapt broad capability to a task&lt;/td&gt;
&lt;td&gt;Fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decide whether a model meets a bar&lt;/td&gt;
&lt;td&gt;Evaluation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Make a model callable&lt;/td&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learn from real usage&lt;/td&gt;
&lt;td&gt;Feedback&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A scenario can require several answers. Long documents do not automatically imply embeddings; they imply chunking first. Search by meaning points to embeddings. Producing a plain-language summary points to a transformer-based language model. Producing an image from text points to a diffusion model.&lt;/p&gt;

&lt;p&gt;The recurring failure mode is choosing “a language model” for the entire scenario. Instead, assign one job to each mechanism and preserve the order in which those jobs occur.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A foundation model is defined by broad pre-training and downstream adaptability, not size alone.&lt;/li&gt;
&lt;li&gt;Tokens are learned sub-word units, and both input and output are counted.&lt;/li&gt;
&lt;li&gt;Chunking divides long material so it can be processed and retrieved with useful context.&lt;/li&gt;
&lt;li&gt;An embedding is a vector whose position encodes meaning; it enables comparison, not reconstruction.&lt;/li&gt;
&lt;li&gt;Keyword search matches strings, while embedding search matches meaning.&lt;/li&gt;
&lt;li&gt;The lifecycle runs from data selection through feedback, with pre-training before fine-tuning.&lt;/li&gt;
&lt;li&gt;Feedback closes the loop by supplying evidence for later adaptation.&lt;/li&gt;
&lt;li&gt;In mixed scenarios, separate preparation, retrieval, generation and lifecycle stages instead of assigning every job to one model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These distinctions settle what tokens, embeddings and lifecycle stages mean and how to recognise them in AIF-C01 scenarios. They do not determine exact token counts, the best chunk size for every document, or whether a deployed model meets a particular quality bar; those answers depend on the tokenizer, the material, the retrieval task and the evaluation criteria.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>certification</category>
    </item>
    <item>
      <title>Give Your Chatbot a Memory in Google Colab Before Your Next AI Interview</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Wed, 15 Jul 2026 13:34:07 +0000</pubDate>
      <link>https://dev.to/confident_prep/give-your-chatbot-a-memory-in-google-colab-before-your-next-ai-interview-4jbb</link>
      <guid>https://dev.to/confident_prep/give-your-chatbot-a-memory-in-google-colab-before-your-next-ai-interview-4jbb</guid>
      <description>&lt;p&gt;You don't need a vector database to understand LLM memory.&lt;/p&gt;

&lt;p&gt;You don't need LangChain.&lt;/p&gt;

&lt;p&gt;You don't need an API key.&lt;/p&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google Colab&lt;/li&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;A conversation that keeps growing&lt;/li&gt;
&lt;li&gt;About 15 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In &lt;a href="https://dev.to/pandeyc005/build-a-rag-system-in-google-colab-before-your-next-ai-interview-3jd3"&gt;an earlier article&lt;/a&gt;, we built the retrieval half of RAG — chunking, embeddings, cosine similarity. This one builds the other half every multi-turn LLM app needs: memory. It's the exact mechanism behind the interview question, "how do you manage memory in an LLM application when conversations get long?"&lt;/p&gt;

&lt;p&gt;Let's build it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Are We Building?
&lt;/h2&gt;

&lt;p&gt;An LLM has no memory of its own. Every call only sees what you send it — the context window. Left alone, a growing conversation does this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Turn 1 → Turn 2 → Turn 3 → ... → Turn 50
              ↓
   Send the full history every time
              ↓
   Context window fills, cost climbs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We'll build the fix: a token counter, a summarization trigger, and a &lt;code&gt;remember()&lt;/code&gt; function that keeps a conversation coherent without sending everything, every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: A Conversation That Won't Stop Growing
&lt;/h2&gt;

&lt;p&gt;Open a Colab notebook and simulate 50 turns of chat — no API needed yet.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;approx_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;0.75&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# a token is roughly 3/4 of a word
&lt;/span&gt;
&lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;51&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Question &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; about our product.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Answer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;, referencing prior context.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Total turns stored:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 2: The Cost of "Just Send Everything"
&lt;/h2&gt;

&lt;p&gt;The simplest memory strategy — an in-context buffer — resends the full history on every call. Let's measure what that actually costs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cost_of_turn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_turns&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;sent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt; &lt;span class="n"&gt;n_turns&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;approx_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;turn_2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cost_of_turn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;turn_50&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cost_of_turn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Turn 2 cost:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;turn_2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Turn 50 cost:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;turn_50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Turn 50 is&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;turn_50&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;turn_2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x more expensive than turn 2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it. Turn 50 costs roughly &lt;strong&gt;25x&lt;/strong&gt; what turn 2 costs — because turn 50 resends 49 prior turns as input. This is the token cost blowout, and it's a cost problem, not just a UX one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Add a Summarization Trigger
&lt;/h2&gt;

&lt;p&gt;Once history crosses a token threshold, compress the old turns into a summary and drop the rest.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SUMMARY_THRESHOLD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;  &lt;span class="c1"&gt;# tokens -- small on purpose, so it triggers a few times in 50 turns
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;summarize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;turns&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;goal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;turns&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;facts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;turns&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;plan&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Goal: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | Facts: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;facts&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;none stated&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; | Compressed &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;turns&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; turns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;remember&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_msg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;convo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_msg&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
    &lt;span class="n"&gt;full_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;approx_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;full_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;SUMMARY_THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;summarize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;convo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A real summary can't just be "the last N words" — that's how you lose facts that fell in the &lt;em&gt;middle&lt;/em&gt; of the conversation. LLMs already attend poorly to the middle of a long context (a property called "lost in the middle"); a lossy summary makes it worse. A summary that survives should always carry four things: the user's goal, key decisions made, open items, and any persistent facts (name, plan tier, stated preferences). Drop any of those and the assistant re-asks a question it already answered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Watch the Cost Curve Bend
&lt;/h2&gt;

&lt;p&gt;Re-run the growing conversation, but through &lt;code&gt;remember()&lt;/code&gt; this time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
&lt;span class="n"&gt;costs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;51&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;remember&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Question &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Answer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="n"&gt;sent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;convo&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;costs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;approx_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sent&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Turn 2 cost:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;costs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Turn 50 cost:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;costs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;49&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare &lt;code&gt;costs[49]&lt;/code&gt; here to &lt;code&gt;turn_50&lt;/code&gt; from Step 2 — roughly 28 tokens against roughly 666. It doesn't keep climbing — it stays bounded, because old turns keep getting folded into a fixed-size summary instead of piling up forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make It Reusable
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;remember()&lt;/code&gt; is already the reusable piece. In a real app, &lt;code&gt;summarize()&lt;/code&gt; is one extra LLM call instead of the string-matching stub above — same shape, smarter compression. Everything else — the threshold check, the token counter, the trigger — stays exactly as it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five Colab Experiments to Try
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Set &lt;code&gt;SUMMARY_THRESHOLD = 100&lt;/code&gt; and then &lt;code&gt;1000&lt;/code&gt;. Watch summarization trigger less and less — this threshold is a real production tuning knob.&lt;/li&gt;
&lt;li&gt;Put a fact in turn 1 ("I'm on the Enterprise plan"), then run 30 turns and print &lt;code&gt;summary&lt;/code&gt; after each compression. It survives the &lt;em&gt;first&lt;/em&gt; compression — then quietly disappears on the second one, because &lt;code&gt;summarize()&lt;/code&gt; only reads the current window, not the running summary. Watch it happen.&lt;/li&gt;
&lt;li&gt;Now fix it: change &lt;code&gt;summarize()&lt;/code&gt; to also accept the current &lt;code&gt;summary&lt;/code&gt; string and fold its contents into the new one, instead of overwriting it. Rerun the same 30-turn test — the fact should survive indefinitely this time.&lt;/li&gt;
&lt;li&gt;Print &lt;code&gt;costs&lt;/code&gt; from Step 4 next to &lt;code&gt;[cost_of_turn(n) for n in range(1, 51)]&lt;/code&gt; from Step 2. Chart both — that's the cost-vs-turn graph an interviewer is picturing when they ask about scaling memory.&lt;/li&gt;
&lt;li&gt;Add a &lt;code&gt;user_memory.json&lt;/code&gt; file that saves one fact outside &lt;code&gt;history&lt;/code&gt;/&lt;code&gt;summary&lt;/code&gt; entirely, and load it into a fresh &lt;code&gt;convo&lt;/code&gt;/&lt;code&gt;summary&lt;/code&gt; pair on a new run — that's semantic memory (persists across sessions), as opposed to episodic memory (resets with the conversation).&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Interview Questions Hidden Inside This Notebook
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why not just use a bigger context window?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Cost and attention. A 200K-token context costs far more per call than a 4K one when full, and "lost in the middle" gets worse, not better, as the window grows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between an in-context buffer and summarization memory?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A buffer resends everything and is simplest but unbounded in cost. Summarization compresses old turns into a fixed-size summary once a token threshold is crossed, keeping cost roughly flat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is "lost in the middle"?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Even when a full context fits the window, models recall the beginning and end more reliably than the middle. Summaries have to explicitly preserve key facts — not just compress chronologically — or those facts vanish first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is memory design a cost decision or a UX decision?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both, but cost drives the threshold. Turn 50 on a raw buffer can cost 25x turn 2 — that's a unit-economics problem before it's ever a UX one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between episodic and semantic memory?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Episodic memory is what happened this conversation and resets at session end. Semantic memory is persistent facts about the user (plan tier, preferences) that survive across sessions and get injected as context at the start of a new one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Memory Manager Is the Easy Part
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;remember()&lt;/code&gt; function above is maybe fifteen lines. A production memory layer also has to handle:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context poisoning — a bad instruction from turn 1 surviving into every later summary&lt;/li&gt;
&lt;li&gt;composite retrieval queries — for external/vector memory, a follow-up like "show me another one" has no retrieval signal on its own, so the query needs the summary plus the last couple of turns, not just the latest message&lt;/li&gt;
&lt;li&gt;validating that summaries never overwrite system rules with user-turn content&lt;/li&gt;
&lt;li&gt;knowing when to graduate: buffer for short sessions, summarization once you cross a few thousand tokens, external retrieval only once sessions span hours or multiple visits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Interviewers ask about these because they want to know you've operated a system like this, not just described one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to Go Deeper?
&lt;/h2&gt;

&lt;p&gt;If running this left you asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do I pick a good summarization threshold for my actual traffic?&lt;/p&gt;

&lt;p&gt;When do I need external memory instead of summarization?&lt;/p&gt;

&lt;p&gt;How do I test that my summary isn't silently dropping facts?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are exactly the questions the full session covers — plus the 5-beat interview answer for "how do you manage memory in an LLM application":&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/4-context-and-memory-management/" rel="noopener noreferrer"&gt;https://confidentprep.com/courses/ai-ml-for-interview/4-context-and-memory-management/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And if you'd rather work through this with other developers and ask questions live, join one of the upcoming live sessions:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/live/" rel="noopener noreferrer"&gt;https://confidentprep.com/live/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Build the small version first. Watch it break at turn 50. Then learn how to explain why — that's a better way to prepare for an AI interview than memorizing another list of memory-management definitions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>tutorial</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>AI Is Creating More Opportunities Than We Realize</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Tue, 14 Jul 2026 16:35:44 +0000</pubDate>
      <link>https://dev.to/confident_prep/ai-is-creating-more-opportunities-than-we-realize-10cg</link>
      <guid>https://dev.to/confident_prep/ai-is-creating-more-opportunities-than-we-realize-10cg</guid>
      <description>&lt;p&gt;The biggest fear around AI is simple: &lt;strong&gt;If two people with AI can do the work of ten people, what happens to the other eight?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is a valid concern. AI will automate tasks. Some teams will shrink. Some skills will lose value.&lt;/p&gt;

&lt;p&gt;But almost every discussion about AI replacing jobs makes one major assumption:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The amount of software the world wants to build will remain constant.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I don't think it will.&lt;/p&gt;

&lt;p&gt;When building becomes dramatically cheaper, we don't just build the same things with fewer people. &lt;strong&gt;We build more things.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. One Startup Needed 10 People. Now One Founder Can Build Four Startups.
&lt;/h2&gt;

&lt;p&gt;Imagine that building a software company previously required ten people. Today, AI coding agents, cloud platforms, automation, and AI-powered support might allow two or three people to build the same product.&lt;/p&gt;

&lt;p&gt;The obvious conclusion is: &lt;strong&gt;eight jobs disappeared.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But consider another possibility. What if the founder who could previously afford to build &lt;strong&gt;one startup can now launch four products in parallel?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Before:&lt;/strong&gt; 1 startup × 10 people = 10 opportunities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;After:&lt;/strong&gt; 4 startups × 2–3 people = 8–12 opportunities&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The individual companies became smaller, but the number of companies, products, and experiments increased.&lt;/p&gt;

&lt;p&gt;We are already seeing early signs of this. In 2026, the startup JustPaid reportedly created a team of seven AI coding agents using OpenClaw and Claude Code. In one month, those agents built 10 major features.&lt;/p&gt;

&lt;p&gt;Instead of eliminating every human role, the company redirected people toward higher-priority customer work and even hired a new developer who was trained largely by the AI agents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI doesn't only reduce the number of people required to build something. It increases the number of things people can afford to build.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. We Have Seen This Before With Cloud Computing
&lt;/h2&gt;

&lt;p&gt;Before cloud computing, companies needed people to buy and install servers, manage operating systems, configure networks, maintain storage and backups, provision infrastructure, and operate physical data centers.&lt;/p&gt;

&lt;p&gt;Cloud computing automated or eliminated many of these responsibilities.&lt;/p&gt;

&lt;p&gt;But the technology industry did not disappear. Instead, we created new careers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloud engineers and solution architects&lt;/li&gt;
&lt;li&gt;DevOps, platform, and Site Reliability Engineers&lt;/li&gt;
&lt;li&gt;Cloud security specialists&lt;/li&gt;
&lt;li&gt;Infrastructure automation and Kubernetes engineers&lt;/li&gt;
&lt;li&gt;FinOps engineers and cloud consultants&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cloud computing removed work, but it also made technology cheaper and accessible to millions of additional companies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI could create the same transformation—much faster.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Millions of AI Agents Will Need Humans Around Them
&lt;/h2&gt;

&lt;p&gt;We are moving from AI systems that answer questions to AI agents that take actions.&lt;/p&gt;

&lt;p&gt;Agents can deploy infrastructure, modify production systems, process customer requests, approve transactions, access company data, and create cloud resources.&lt;/p&gt;

&lt;p&gt;Now imagine an AI cloud agent making the wrong decision. It could delete production infrastructure or create millions of dollars in cloud costs.&lt;/p&gt;

&lt;p&gt;Or imagine a support agent processing one million tickets with a 2% serious error rate. That is &lt;strong&gt;20,000 incorrect decisions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Someone still has to answer critical questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the agent safe enough to deploy?&lt;/li&gt;
&lt;li&gt;Which actions can it perform autonomously?&lt;/li&gt;
&lt;li&gt;Which actions require human approval?&lt;/li&gt;
&lt;li&gt;How much money or infrastructure can it control?&lt;/li&gt;
&lt;li&gt;Who monitors its behavior?&lt;/li&gt;
&lt;li&gt;Who investigates failures?&lt;/li&gt;
&lt;li&gt;Who stops it when something goes wrong?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The more autonomous AI becomes, the more valuable human judgment becomes.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The AI Evaluation and Security Economy Is Just Beginning
&lt;/h2&gt;

&lt;p&gt;Traditional software testing will not be enough. AI systems are probabilistic, and agents increasingly have access to real tools, money, infrastructure, and data.&lt;/p&gt;

&lt;p&gt;We will need people and platforms focused on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI evaluations and continuous testing&lt;/li&gt;
&lt;li&gt;Agent observability and monitoring&lt;/li&gt;
&lt;li&gt;Guardrails and permission boundaries&lt;/li&gt;
&lt;li&gt;Human-in-the-loop approval systems&lt;/li&gt;
&lt;li&gt;AI security and prompt-injection defense&lt;/li&gt;
&lt;li&gt;Cost controls and anomaly detection&lt;/li&gt;
&lt;li&gt;Audit trails, compliance, and governance&lt;/li&gt;
&lt;li&gt;Red teaming and failure investigation&lt;/li&gt;
&lt;li&gt;Rollback and recovery mechanisms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cloud computing created cloud security. APIs created API security. Containers created container security. AI agents will create entirely new security, governance, and operational problems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Automation does not eliminate responsibility. It increases the scale at which mistakes can happen.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Solving these problems will create companies, products, and careers that barely exist today.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The Biggest Advantage Will Belong to People Who Keep Learning
&lt;/h2&gt;

&lt;p&gt;AI is moving incredibly fast. Skills that are valuable today may become automated tomorrow.&lt;/p&gt;

&lt;p&gt;That creates an unusual opportunity because nobody has twenty years of experience building production AI agents, agent observability platforms, AI evaluation systems, or AI-native software companies.&lt;/p&gt;

&lt;p&gt;The playing field has partially reset.&lt;/p&gt;

&lt;p&gt;Young professionals can enter emerging fields before established career paths exist. Experienced professionals can combine decades of engineering and business judgment with powerful AI tools.&lt;/p&gt;

&lt;p&gt;The advantage will not automatically belong to the youngest or most experienced person.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It will belong to the person willing to keep learning.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Don't Compete With AI. Become Excellent at Using It.
&lt;/h2&gt;

&lt;p&gt;If AI becomes excellent at generating repetitive code, don't build your entire career around writing repetitive code. Move one level higher.&lt;/p&gt;

&lt;p&gt;Learn how to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Build real systems using AI&lt;/li&gt;
&lt;li&gt;Evaluate and monitor AI agents&lt;/li&gt;
&lt;li&gt;Secure autonomous systems&lt;/li&gt;
&lt;li&gt;Design human-in-the-loop workflows&lt;/li&gt;
&lt;li&gt;Integrate AI with real businesses&lt;/li&gt;
&lt;li&gt;Identify valuable problems worth solving&lt;/li&gt;
&lt;li&gt;Use AI to launch products that were previously too expensive to build&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The most valuable technology professional of the AI era may not be the person who writes the most code.&lt;/p&gt;

&lt;p&gt;It may be the person who knows:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What should we build? How can AI help us build it? How do we know it actually works? And how do we operate it safely at scale?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;AI will eliminate some jobs. But it will also make thousands of new products economically possible.&lt;/p&gt;

&lt;p&gt;Those products will need to be built, integrated, evaluated, secured, monitored, governed, improved, and turned into businesses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The amount of opportunity in the world is not fixed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI is expanding what individuals and small teams can attempt to build. The biggest opportunity is not in competing with the tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is in becoming exceptionally good at using it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to Start Learning AI?
&lt;/h2&gt;

&lt;p&gt;The best way to prepare for the AI era is to start building, experimenting, and learning how these systems actually work.&lt;/p&gt;

&lt;p&gt;If you prefer learning interactively, &lt;strong&gt;join the upcoming live AI sessions:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://confidentprep.com/live/" rel="noopener noreferrer"&gt;https://confidentprep.com/live/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or learn at your own pace with &lt;strong&gt;self-paced AI courses and practical assignments:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://confidentprep.com/courses/" rel="noopener noreferrer"&gt;https://confidentprep.com/courses/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't wait until AI changes your role to start learning AI. Start learning how to use it now.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>discuss</category>
      <category>interview</category>
    </item>
    <item>
      <title>This RAG Interview Question Looks Simple. Many Candidates Still Get It Wrong.</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Tue, 14 Jul 2026 16:18:15 +0000</pubDate>
      <link>https://dev.to/confident_prep/this-rag-interview-question-looks-simple-many-candidates-still-get-it-wrong-1d14</link>
      <guid>https://dev.to/confident_prep/this-rag-interview-question-looks-simple-many-candidates-still-get-it-wrong-1d14</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;You have converted 10,000 documents into embeddings and stored them in a vector database. A user asks a question. Explain exactly how the system finds the most relevant chunks to send to the LLM.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Could you answer this clearly in an interview?&lt;/p&gt;

&lt;p&gt;Not just say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We perform semantic search."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But explain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what happens to the user's question&lt;/li&gt;
&lt;li&gt;why the question must be converted into an embedding&lt;/li&gt;
&lt;li&gt;how it is compared with stored document embeddings&lt;/li&gt;
&lt;li&gt;what cosine similarity is doing&lt;/li&gt;
&lt;li&gt;how the chunks are ranked&lt;/li&gt;
&lt;li&gt;what Top-K retrieval means&lt;/li&gt;
&lt;li&gt;and what is finally sent to the LLM&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;You could easily answer if you read this article:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/pandeyc005/build-a-rag-system-in-google-colab-before-your-next-ai-interview-3jd3"&gt;Build a RAG System in Google Colab Before Your Next AI Interview&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The article makes you build the entire retrieval flow yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Documents
    ↓
Chunks
    ↓
Embeddings
    ↓
User Question
    ↓
Query Embedding
    ↓
Cosine Similarity
    ↓
Ranked Chunks
    ↓
Top-K Results
    ↓
LLM Prompt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And this is exactly why building a small RAG system before an AI interview is useful.&lt;/p&gt;

&lt;p&gt;Many candidates know the architecture diagram.&lt;/p&gt;

&lt;p&gt;Far fewer can explain what actually happens between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Question → Vector Search → Retrieved Context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once you implement it yourself, questions about embeddings, similarity search, Top-K retrieval, chunking, and context construction become much easier to answer.&lt;/p&gt;

&lt;p&gt;But here is the follow-up question interviewers may ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If the most relevant chunk is ranked #6, but your system retrieves only Top-5 chunks, what happens? How would you improve the system?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now the discussion moves into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;choosing &lt;code&gt;top_k&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;retrieval recall&lt;/li&gt;
&lt;li&gt;reranking&lt;/li&gt;
&lt;li&gt;hybrid search&lt;/li&gt;
&lt;li&gt;better chunking&lt;/li&gt;
&lt;li&gt;metadata filtering&lt;/li&gt;
&lt;li&gt;retrieval evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is where understanding RAG becomes more important than simply knowing how to call a vector database.&lt;/p&gt;

&lt;p&gt;If you're preparing for AI/ML interviews, you can continue learning about embeddings, vector search, and RAG here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/3-embeddings-and-rag/" rel="noopener noreferrer"&gt;Embeddings and RAG Interview Preparation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And if you'd like to practice these concepts through live discussions and interview questions:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/live/" rel="noopener noreferrer"&gt;Attend an Upcoming Free Live Session&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>interview</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Build a RAG System in Google Colab Before Your Next AI Interview</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Sun, 12 Jul 2026 17:35:47 +0000</pubDate>
      <link>https://dev.to/confident_prep/build-a-rag-system-in-google-colab-before-your-next-ai-interview-3jd3</link>
      <guid>https://dev.to/confident_prep/build-a-rag-system-in-google-colab-before-your-next-ai-interview-3jd3</guid>
      <description>&lt;p&gt;You don't need to study machine learning for six months to understand RAG.&lt;/p&gt;

&lt;p&gt;You don't need Kubernetes.&lt;/p&gt;

&lt;p&gt;You don't need a vector database.&lt;/p&gt;

&lt;p&gt;You don't even need an OpenAI API key.&lt;/p&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google Colab&lt;/li&gt;
&lt;li&gt;Python&lt;/li&gt;
&lt;li&gt;A few paragraphs of text&lt;/li&gt;
&lt;li&gt;About 20 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By the end of this article, you will have built the core retrieval mechanism behind a RAG system.&lt;/p&gt;

&lt;p&gt;More importantly, you will understand what is actually happening when someone says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We convert documents into embeddings, store them in a vector database, retrieve relevant chunks, and send them to an LLM."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Let's build it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Are We Building?
&lt;/h2&gt;

&lt;p&gt;Imagine that we have a small knowledge base about AWS services.&lt;/p&gt;

&lt;p&gt;A user asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which AWS service should I use to decouple applications?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Our system will:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Break documents into chunks.&lt;/li&gt;
&lt;li&gt;Convert the chunks into embeddings.&lt;/li&gt;
&lt;li&gt;Convert the question into an embedding.&lt;/li&gt;
&lt;li&gt;Compare the question with every chunk.&lt;/li&gt;
&lt;li&gt;Retrieve the most relevant chunks.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the &lt;strong&gt;retrieval&lt;/strong&gt; part of Retrieval-Augmented Generation.&lt;/p&gt;

&lt;p&gt;The architecture looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Documents
    ↓
Chunking
    ↓
Embeddings
    ↓
Vector Store

User Question
    ↓
Question Embedding
    ↓
Similarity Search
    ↓
Relevant Chunks
    ↓
LLM
    ↓
Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a production application, you might use Pinecone, OpenSearch, pgvector, Qdrant, or another vector database.&lt;/p&gt;

&lt;p&gt;For learning, we don't need any of them.&lt;/p&gt;

&lt;p&gt;We will store our embeddings in memory and use cosine similarity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Open Google Colab
&lt;/h2&gt;

&lt;p&gt;Create a new Google Colab notebook.&lt;/p&gt;

&lt;p&gt;Run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="err"&gt;!&lt;/span&gt;&lt;span class="n"&gt;pip&lt;/span&gt; &lt;span class="n"&gt;install&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;q&lt;/span&gt; &lt;span class="n"&gt;sentence&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;transformers&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We will use the &lt;code&gt;sentence-transformers&lt;/code&gt; library to generate embeddings.&lt;/p&gt;

&lt;p&gt;No API key is required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Create Our Knowledge Base
&lt;/h2&gt;

&lt;p&gt;Let's create a tiny collection of documents.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Amazon S3 is an object storage service designed for storing
    and retrieving files. It provides high durability and is
    commonly used for backups, static websites, data lakes,
    and application assets.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Amazon SQS is a managed message queue service.

    It allows applications to communicate asynchronously.

    Producers send messages to a queue and consumers process
    those messages independently.

    SQS is commonly used to decouple distributed applications.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    AWS Lambda is a serverless compute service.

    Developers upload code and AWS executes the code in response
    to events.

    Lambda automatically manages servers and scales applications
    based on incoming requests.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;

    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Amazon DynamoDB is a managed NoSQL database.

    It provides low-latency access to data and automatically
    scales to handle large workloads.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Obviously, real RAG systems contain thousands or millions of documents.&lt;/p&gt;

&lt;p&gt;But the mechanism is exactly the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Chunk the Documents
&lt;/h2&gt;

&lt;p&gt;Why do we need chunks?&lt;/p&gt;

&lt;p&gt;Because embedding an entire book or a 200-page PDF as one vector would produce a poor representation for individual questions.&lt;/p&gt;

&lt;p&gt;Instead, RAG systems divide documents into smaller pieces.&lt;/p&gt;

&lt;p&gt;Let's write a very simple chunker.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;250&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;words&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;words&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;words&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;


&lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;


&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Number of chunks:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;CHUNK &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Production systems usually use more sophisticated strategies.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;token-based chunking&lt;/li&gt;
&lt;li&gt;overlapping chunks&lt;/li&gt;
&lt;li&gt;recursive text splitting&lt;/li&gt;
&lt;li&gt;semantic chunking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But our goal is to understand the mechanism first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Generate Embeddings
&lt;/h2&gt;

&lt;p&gt;Now we need to convert text into vectors.&lt;/p&gt;

&lt;p&gt;We will use a small embedding model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now generate embeddings for our chunks.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chunk_embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk_embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see something similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(4, 384)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What does that mean?&lt;/p&gt;

&lt;p&gt;We have four chunks.&lt;/p&gt;

&lt;p&gt;Each chunk has been converted into a vector containing 384 numbers.&lt;/p&gt;

&lt;p&gt;Something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[0.023, -0.041, 0.087, ...]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Humans cannot interpret these numbers directly.&lt;/p&gt;

&lt;p&gt;But mathematically, texts with similar meanings tend to have vectors that are closer together.&lt;/p&gt;

&lt;p&gt;That is what makes semantic search possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Ask a Question
&lt;/h2&gt;

&lt;p&gt;Let's ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which AWS service can help decouple applications?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Convert the question into an embedding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;question_embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now we have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Documents → Vectors

Question → Vector
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We need to find which document vectors are closest to the question vector.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Perform Similarity Search
&lt;/h2&gt;

&lt;p&gt;We will use cosine similarity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.metrics.pairwise&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cosine_similarity&lt;/span&gt;

&lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;question_embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;chunk_embeddings&lt;/span&gt;
&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's inspect the results.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;][:&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see the SQS document receive the highest similarity score.&lt;/p&gt;

&lt;p&gt;Now retrieve the best chunk.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;best_chunk_index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;retrieved_chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;best_chunk_index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retrieved_chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Congratulations.&lt;/p&gt;

&lt;p&gt;You just built semantic retrieval.&lt;/p&gt;

&lt;p&gt;This is the foundation of a RAG system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let's Make It Reusable
&lt;/h2&gt;

&lt;p&gt;Let's wrap everything into a function.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

    &lt;span class="n"&gt;question_embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

    &lt;span class="n"&gt;scores&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cosine_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;question_embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;chunk_embeddings&lt;/span&gt;
    &lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;top_indices&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;argsort&lt;/span&gt;&lt;span class="p"&gt;()[::&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][:&lt;/span&gt;&lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;top_indices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;})&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now try asking different questions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;questions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Where should I store application files?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How can I run code without managing servers?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which database provides low latency access?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How can microservices communicate asynchronously?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;questions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;QUESTION:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;score&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][:&lt;/span&gt;&lt;span class="mi"&gt;150&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You now have a tiny semantic search engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  But Where Is the LLM?
&lt;/h2&gt;

&lt;p&gt;This is where many developers misunderstand RAG.&lt;/p&gt;

&lt;p&gt;The vector database does not answer the question.&lt;/p&gt;

&lt;p&gt;The embedding model does not answer the question.&lt;/p&gt;

&lt;p&gt;The retrieval system finds relevant information.&lt;/p&gt;

&lt;p&gt;The retrieved information is then placed into the prompt sent to the LLM.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;retrieved_chunks&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retrieved_chunks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Answer the question using only the context below.

    CONTEXT:

    &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

    QUESTION:

    &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's generate the prompt.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which AWS service should I use to decouple applications?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;retrieved_chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;retrieved_chunks&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prompt could now be sent to an LLM.&lt;/p&gt;

&lt;p&gt;That completes the RAG pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Documents
    ↓
Chunks
    ↓
Embeddings
    ↓
Vector Search
    ↓
Relevant Context
    ↓
Prompt
    ↓
LLM
    ↓
Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Five Experiments You Should Try
&lt;/h2&gt;

&lt;p&gt;Don't stop after running the notebook.&lt;/p&gt;

&lt;p&gt;Change it.&lt;/p&gt;

&lt;p&gt;First, add documents that use similar terminology.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SQS decouples applications.

EventBridge connects applications using events.

SNS distributes messages to multiple subscribers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Does the correct document still rank first?&lt;/p&gt;

&lt;p&gt;Second, change the chunk size.&lt;/p&gt;

&lt;p&gt;Try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chunk_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chunk_size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What happens to retrieval quality?&lt;/p&gt;

&lt;p&gt;Third, retrieve more documents.&lt;/p&gt;

&lt;p&gt;Change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;top_k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;top_k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Would sending more context always produce a better answer?&lt;/p&gt;

&lt;p&gt;Fourth, add irrelevant documents.&lt;/p&gt;

&lt;p&gt;Does retrieval quality change?&lt;/p&gt;

&lt;p&gt;Fifth, ask ambiguous questions.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which service should I use for messaging?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you are starting to encounter the problems that real RAG systems must solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Interview Questions Hidden Inside This Project
&lt;/h2&gt;

&lt;p&gt;If you understand the notebook above, you should be able to discuss several common AI interview questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why do RAG systems chunk documents?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because retrieval usually needs to identify specific passages rather than entire documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is an embedding?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A numerical representation of data that allows semantic relationships to be compared mathematically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why use cosine similarity?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because it measures the direction between vectors and is commonly used to compare embedding similarity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a vector database generate answers?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;It stores vectors and helps retrieve relevant information.&lt;/p&gt;

&lt;p&gt;The LLM generates the answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does retrieving more chunks always improve the answer?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;More context can increase cost, introduce irrelevant information, and make it harder for the model to identify the useful evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when the correct document isn't retrieved?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The LLM never receives the information required to answer correctly.&lt;/p&gt;

&lt;p&gt;This is why evaluating retrieval quality is one of the most important parts of building a production RAG system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 20-Line Demo Is the Easy Part
&lt;/h2&gt;

&lt;p&gt;You can build a RAG demo in twenty minutes.&lt;/p&gt;

&lt;p&gt;Production RAG systems are harder.&lt;/p&gt;

&lt;p&gt;You have to make decisions about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;chunk size and overlap&lt;/li&gt;
&lt;li&gt;embedding models&lt;/li&gt;
&lt;li&gt;metadata filtering&lt;/li&gt;
&lt;li&gt;top-k retrieval&lt;/li&gt;
&lt;li&gt;similarity thresholds&lt;/li&gt;
&lt;li&gt;hybrid search&lt;/li&gt;
&lt;li&gt;reranking&lt;/li&gt;
&lt;li&gt;context-window management&lt;/li&gt;
&lt;li&gt;hallucination control&lt;/li&gt;
&lt;li&gt;retrieval evaluation&lt;/li&gt;
&lt;li&gt;answer evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And interviewers increasingly ask about these trade-offs.&lt;/p&gt;

&lt;p&gt;Not because they expect you to memorize definitions.&lt;/p&gt;

&lt;p&gt;They want to know whether you understand how the system behaves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to Go Deeper?
&lt;/h2&gt;

&lt;p&gt;If you ran this notebook and found yourself asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How should I choose chunk size?&lt;/p&gt;

&lt;p&gt;When should I use a vector database?&lt;/p&gt;

&lt;p&gt;How do I know whether retrieval is actually working?&lt;/p&gt;

&lt;p&gt;What is the difference between semantic search, hybrid search, and reranking?&lt;/p&gt;

&lt;p&gt;How would I design this system for thousands or millions of documents?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are exactly the questions worth exploring next.&lt;/p&gt;

&lt;p&gt;I'm covering embeddings, vector search, RAG architecture, and the engineering decisions behind production AI systems on ConfidentPrep.&lt;/p&gt;

&lt;p&gt;You can explore the deeper RAG learning material here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/courses/ai-ml-for-interview/3-embeddings-and-rag/" rel="noopener noreferrer"&gt;https://confidentprep.com/courses/ai-ml-for-interview/3-embeddings-and-rag/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And if you'd rather learn by discussing these concepts live, asking questions, and working through interview scenarios with other developers, join one of the upcoming live sessions:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://confidentprep.com/live/" rel="noopener noreferrer"&gt;https://confidentprep.com/live/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Build the small version first.&lt;/p&gt;

&lt;p&gt;Understand why it works.&lt;/p&gt;

&lt;p&gt;Then learn how to explain the trade-offs.&lt;/p&gt;

&lt;p&gt;That's a much better way to prepare for an AI interview than memorizing another list of 100 RAG questions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>beginners</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Your New Prompt 'Feels' Better. That's Not an Eval.</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Sat, 11 Jul 2026 14:59:57 +0000</pubDate>
      <link>https://dev.to/confident_prep/your-new-prompt-feels-better-thats-not-an-eval-3gan</link>
      <guid>https://dev.to/confident_prep/your-new-prompt-feels-better-thats-not-an-eval-3gan</guid>
      <description>&lt;p&gt;You tweak the prompt. Run it against the three examples you always use to sanity-check. It looks better. Ship it.&lt;/p&gt;

&lt;p&gt;That's not evaluation. That's vibes with extra steps — and it's the exact habit interviewers are trained to catch with one question.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🎯 &lt;strong&gt;TL;DR:&lt;/strong&gt; One interview question, one real framework: don't trust "it looks better on my usual examples" — build a small eval set, score it the same way every time, and re-run it on every future change. Read the framework, then use it as the literal answer next time someone asks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  🎭 The question: "How do you know your new prompt is actually better — not just better on the five examples you tried by hand?"
&lt;/h2&gt;

&lt;p&gt;This is one of the fastest ways an interviewer separates "I iterate on prompts" from "I've actually shipped prompt changes to production." Everyone iterates. Almost nobody evaluates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to answer it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Name the trap first, out loud: eyeballing a handful of hand-picked examples is confirmation bias, not evaluation. You'll always find a few cases where the new version looks better — that's not evidence, that's cherry-picking with good intentions.&lt;/li&gt;
&lt;li&gt;  Describe a minimal eval set: 20-50 real or realistic input/expected-output pairs, run automatically against both the old and new prompt, scored the same way every time — not re-read by eye each round.&lt;/li&gt;
&lt;li&gt;  Say what "scored" actually means, because this is where most candidates go vague: exact match works for structured output (JSON fields, classifications); for open-ended text, you need a similarity or grounding score against a reference answer.&lt;/li&gt;
&lt;li&gt;  Close with the regression angle: the eval set isn't a one-time gate before this ship — it's something you re-run on every future prompt or retrieval change, so a "small tweak" three weeks from now doesn't quietly break something the old version got right.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🕳️ Why "it looked better to me" is losing you the room
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable part: this question isn't really testing whether you know what an eval set is. It's testing whether you've been burned by &lt;em&gt;not&lt;/em&gt; having one — shipped a "better" prompt that regressed silently, found out from a user complaint instead of a dashboard. Candidates who've lived that answer differently than candidates who haven't, and interviewers can usually tell within one follow-up question.&lt;/p&gt;

&lt;p&gt;Somewhere in your interview pool right now, another candidate already built that eval habit into their process — not because a course told them to, but because something broke on them once. That's the gap this question is actually measuring.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧 What the eval set looks like when it's not hypothetical
&lt;/h2&gt;

&lt;p&gt;Skip the abstraction. In practice it's a spreadsheet or a JSON file with columns: input, expected/reference answer, prompt-A output and score, prompt-B output and score. Twenty rows is enough to start. The moment you can point to a number instead of a feeling, you've already answered this question better than most candidates do.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔬 Next: the 12 lines of code that do the scoring
&lt;/h2&gt;

&lt;p&gt;Saying "similarity score" in an interview is good. Being able to sketch the actual code is better — and next week's post is exactly that: a Colab-ready snippet that scores whether a generated answer is actually grounded in retrieved context, the same technique behind the eval-set scoring above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The takeaway: "it looks better" is an opinion. A score you can re-run is an answer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;[See the paths that build that judgment →]&lt;br&gt;
&lt;a href="https://confidentprep.com/paths?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=launch-echo-general-article" rel="noopener noreferrer"&gt;https://confidentprep.com/paths&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;🔴 Preparing for AI interviews? Don’t prepare alone. Join a live session where we break down real interview questions like this one, discuss how strong candidates structure their answers, and work through the follow-up questions interviewers actually ask. Bring your questions, test your answers, and learn live with other candidates preparing for AI roles.&lt;/p&gt;

&lt;p&gt;👉 Join the next live AI interview prep session → &lt;br&gt;
&lt;a href="https://confidentprep.com/live/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=launch-echo-general-article" rel="noopener noreferrer"&gt;https://confidentprep.com/live/&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;More from &lt;a href="https://confidentprep.com?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=launch-echo-general-article" rel="noopener noreferrer"&gt;https://confidentprep.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>career</category>
      <category>interview</category>
    </item>
    <item>
      <title>Stop Learning Machine Learning Before GenAI 🤖</title>
      <dc:creator>confident_prep</dc:creator>
      <pubDate>Sat, 11 Jul 2026 04:25:49 +0000</pubDate>
      <link>https://dev.to/confident_prep/stop-learning-machine-learning-before-genai-22gm</link>
      <guid>https://dev.to/confident_prep/stop-learning-machine-learning-before-genai-22gm</guid>
      <description>&lt;p&gt;Yes, you read that right.&lt;/p&gt;

&lt;p&gt;If your goal is to understand Generative AI, build LLM-powered applications, or prepare for a GenAI interview, &lt;strong&gt;you don't need to finish learning Machine Learning first.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yet many developers get stuck here.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Should I learn statistics first?&lt;br&gt;
Then Machine Learning?&lt;br&gt;
Then Deep Learning?&lt;br&gt;
Then neural networks?&lt;br&gt;
Then transformers?&lt;br&gt;
And finally Generative AI?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This roadmap may make sense if your goal is to become an ML Engineer, Data Scientist, or AI researcher.&lt;/p&gt;

&lt;p&gt;But for many software developers and technology professionals, it creates an unnecessary barrier.&lt;/p&gt;

&lt;p&gt;You keep preparing to start.&lt;/p&gt;

&lt;p&gt;But never actually start.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚧 Don't Let the AI Roadmap Block You
&lt;/h2&gt;

&lt;p&gt;Machine Learning is a large and valuable field.&lt;/p&gt;

&lt;p&gt;But you don't need to master regression, classification algorithms, backpropagation, or the mathematics of neural networks before you can understand how modern GenAI applications work.&lt;/p&gt;

&lt;p&gt;Start with a simpler question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What do I need to understand to build and discuss a GenAI application?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer is much more approachable.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧠 Start With Generative AI Fundamentals
&lt;/h2&gt;

&lt;p&gt;Understand the basic concepts first.&lt;/p&gt;

&lt;p&gt;What is Generative AI?&lt;/p&gt;

&lt;p&gt;What is an LLM?&lt;/p&gt;

&lt;p&gt;What is a prompt?&lt;/p&gt;

&lt;p&gt;What are tokens and context windows?&lt;/p&gt;

&lt;p&gt;Why do LLMs hallucinate?&lt;/p&gt;

&lt;p&gt;How does an application communicate with an LLM?&lt;/p&gt;

&lt;p&gt;What happens when you send a prompt and receive a response?&lt;/p&gt;

&lt;p&gt;You don't need to understand every mathematical detail behind the model.&lt;/p&gt;

&lt;p&gt;But you should be able to explain &lt;strong&gt;what these concepts mean, why they matter, and how they affect real applications.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's a good starting point.&lt;/p&gt;

&lt;h2&gt;
  
  
  🛠️ Then Build Something Small
&lt;/h2&gt;

&lt;p&gt;Call an LLM API.&lt;/p&gt;

&lt;p&gt;Send a prompt.&lt;/p&gt;

&lt;p&gt;Get a response.&lt;/p&gt;

&lt;p&gt;Change the prompt and observe what happens.&lt;/p&gt;

&lt;p&gt;Experiment with model parameters.&lt;/p&gt;

&lt;p&gt;Then build a small application.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A document summarizer&lt;/li&gt;
&lt;li&gt;A question-answering application&lt;/li&gt;
&lt;li&gt;A chatbot&lt;/li&gt;
&lt;li&gt;A structured data extraction tool&lt;/li&gt;
&lt;li&gt;A basic RAG application&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While building, you'll naturally encounter new questions.&lt;/p&gt;

&lt;p&gt;How do I provide my own data to an LLM?&lt;/p&gt;

&lt;p&gt;What are embeddings?&lt;/p&gt;

&lt;p&gt;Why do I need a vector database?&lt;/p&gt;

&lt;p&gt;How should documents be chunked?&lt;/p&gt;

&lt;p&gt;How do I reduce hallucinations?&lt;/p&gt;

&lt;p&gt;How do I evaluate the quality of responses?&lt;/p&gt;

&lt;p&gt;How do I protect an application from prompt injection?&lt;/p&gt;

&lt;p&gt;Now you have a reason to learn these concepts.&lt;/p&gt;

&lt;p&gt;You're learning because you need to solve a problem—not because a massive AI roadmap told you to learn everything first.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎯 Preparing for a GenAI Interview?
&lt;/h2&gt;

&lt;p&gt;The same principle applies.&lt;/p&gt;

&lt;p&gt;Don't wait until you know everything about Machine Learning before preparing.&lt;/p&gt;

&lt;p&gt;Start with the questions that help you understand the GenAI application landscape.&lt;/p&gt;

&lt;p&gt;Can you explain how an LLM-powered application works?&lt;/p&gt;

&lt;p&gt;Can you explain tokens, context windows, and hallucinations?&lt;/p&gt;

&lt;p&gt;Do you understand the difference between prompting, RAG, and fine-tuning?&lt;/p&gt;

&lt;p&gt;Can you explain why an application might use embeddings and a vector database?&lt;/p&gt;

&lt;p&gt;Can you discuss security, cost, latency, evaluation, and reliability?&lt;/p&gt;

&lt;p&gt;Can you describe something you have built, even if it is small?&lt;/p&gt;

&lt;p&gt;These questions give you a practical direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚠️ Does This Mean Machine Learning Is Not Important?
&lt;/h2&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;Machine Learning fundamentals become increasingly important depending on the role you are targeting.&lt;/p&gt;

&lt;p&gt;If you want to train models, work deeply with model architectures, become an ML Engineer, or pursue AI research, you will need stronger foundations in Machine Learning, mathematics, and statistics.&lt;/p&gt;

&lt;p&gt;But that's different from saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Everyone must learn Machine Learning before they can start learning Generative AI.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They don't.&lt;/p&gt;

&lt;p&gt;For many developers, architects, testers, DevOps engineers, and other technology professionals, starting with GenAI applications is a perfectly reasonable path.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌱 Start First. Go Deeper When You Need To.
&lt;/h2&gt;

&lt;p&gt;The AI ecosystem is enormous.&lt;/p&gt;

&lt;p&gt;You can spend months creating the perfect learning roadmap.&lt;/p&gt;

&lt;p&gt;Or you can start.&lt;/p&gt;

&lt;p&gt;Understand what Generative AI is.&lt;/p&gt;

&lt;p&gt;Learn the fundamentals.&lt;/p&gt;

&lt;p&gt;Build something small.&lt;/p&gt;

&lt;p&gt;Prepare for practical interview questions.&lt;/p&gt;

&lt;p&gt;Discover your knowledge gaps.&lt;/p&gt;

&lt;p&gt;Then go deeper.&lt;/p&gt;

&lt;p&gt;If you're preparing for a GenAI interview and don't know where to begin, I've put together a structured guide covering the concepts and questions worth exploring:&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://confidentprep.com/interview/ai/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Start Preparing for Your AI Interview&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Don't let the size of Machine Learning stop you from starting with Generative AI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Start small. Understand the fundamentals. Build something. Then go deeper. 🚀&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;→ For more details, see &lt;a href="https://confidentprep.com/interview/ai/" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;More from &lt;a href="https://confidentprep.com?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=launch-echo-general-stop-learning-machine-learning-before-genai" rel="noopener noreferrer"&gt;https://confidentprep.com&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>genai</category>
      <category>beginners</category>
      <category>career</category>
    </item>
  </channel>
</rss>
