<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: yash-nigam</title>
    <description>The latest articles on DEV Community by yash-nigam (@yashnigam).</description>
    <link>https://dev.to/yashnigam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1034882%2F859c3509-9e87-46d8-ae9a-ba1c65d73b13.png</url>
      <title>DEV Community: yash-nigam</title>
      <link>https://dev.to/yashnigam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yashnigam"/>
    <language>en</language>
    <item>
      <title>LLM Concepts Deep Dive - Conceptual Mastery for Developers | Udemy</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Wed, 16 Sep 2026 12:00:03 +0000</pubDate>
      <link>https://dev.to/yashnigam/llm-concepts-deep-dive-conceptual-mastery-for-developers-udemy-5cla</link>
      <guid>https://dev.to/yashnigam/llm-concepts-deep-dive-conceptual-mastery-for-developers-udemy-5cla</guid>
      <description>&lt;h2&gt;
  
  
  LLM Concepts Deep Dive
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Course: LLM Concepts Deep Dive: Conceptual Mastery for Developers | Udemy Business&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Notes from: Monday, September 14, 2026&lt;/em&gt;&lt;br&gt;
*&lt;a href="https://www.udemy.com/course/llm-concepts-deep-dive/?couponCode=25BBPMXNVD35" rel="noopener noreferrer"&gt;https://www.udemy.com/course/llm-concepts-deep-dive/?couponCode=25BBPMXNVD35&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
What is a Model

&lt;ul&gt;
&lt;li&gt;1.1 What is a Language Model?&lt;/li&gt;
&lt;li&gt;1.2 How Does a Language Model Learn?&lt;/li&gt;
&lt;li&gt;1.3 What Does "Predict the Likelihood of Next Set of Words" Mean?&lt;/li&gt;
&lt;li&gt;1.4 Language Model vs. Generative AI&lt;/li&gt;
&lt;li&gt;1.5 Modern LLMs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Two Types of Language Modeling Tasks

&lt;ul&gt;
&lt;li&gt;2.1 Autoencoding&lt;/li&gt;
&lt;li&gt;2.2 Autoregressive&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
LLM Training Methodology

&lt;ul&gt;
&lt;li&gt;3.1 Learn/Model Language&lt;/li&gt;
&lt;li&gt;3.2 Pre-training&lt;/li&gt;
&lt;li&gt;3.3 Weights&lt;/li&gt;
&lt;li&gt;3.4 Pre-training Alone Is Not Enough&lt;/li&gt;
&lt;li&gt;3.5 Fine-tuning&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Tokens and Embeddings

&lt;ul&gt;
&lt;li&gt;4.1 How LLMs Process Tokens&lt;/li&gt;
&lt;li&gt;4.2 Visualizing Tokenization&lt;/li&gt;
&lt;li&gt;4.3 Why Isn't One Word Always One Token?&lt;/li&gt;
&lt;li&gt;4.4 The Solution&lt;/li&gt;
&lt;li&gt;4.5 Token Decoding&lt;/li&gt;
&lt;li&gt;4.6 What Numbers Are Assigned to Tokens&lt;/li&gt;
&lt;li&gt;4.7 How the Model Learns&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Feature Matrix

&lt;ul&gt;
&lt;li&gt;5.1 Feature Matrix for Language&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Embedding

&lt;ul&gt;
&lt;li&gt;6.1 Understanding Embeddings Using Dimensions&lt;/li&gt;
&lt;li&gt;6.2 Embedding Math&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Tokens and Their Values

&lt;ul&gt;
&lt;li&gt;7.1 Text Similarity&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Transformer Architecture&lt;/li&gt;
&lt;li&gt;
Context Length in LLMs

&lt;ul&gt;
&lt;li&gt;9.1 LLM State&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
What is RAG

&lt;ul&gt;
&lt;li&gt;10.1 Vectors and Chunks&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;RAG Pipeline&lt;/li&gt;
&lt;li&gt;
Vector Databases

&lt;ul&gt;
&lt;li&gt;12.1 How It Works — Ingestion&lt;/li&gt;
&lt;li&gt;12.2 How It Works — Indexing&lt;/li&gt;
&lt;li&gt;12.3 How It Works — Querying&lt;/li&gt;
&lt;li&gt;12.4 Some Vector DBs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is a Model
&lt;/h2&gt;

&lt;p&gt;A model is something that represents, simulates, or predicts something else.&lt;br&gt;
It is a simplified version of a real thing that helps us understand it or predict what might happen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;🏠 Model of a house - Represents what a real house looks like.&lt;/li&gt;
&lt;li&gt;🏙️ Model of a city - Represents the roads, buildings, and layout of a real city.&lt;/li&gt;
&lt;li&gt;🌦️ Weather model - Uses weather data and patterns from the past to predict how the weather may be in the future.&lt;/li&gt;
&lt;li&gt;🗣️ Language model - Learns patterns in language and predicts what words or tokens are likely to come next.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  1.1 What is a Language Model?
&lt;/h3&gt;

&lt;p&gt;A language model is a model that learns how language works, its patterns from large amounts of text, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which words commonly appear together.&lt;/li&gt;
&lt;li&gt;How sentences are structured.&lt;/li&gt;
&lt;li&gt;How people communicate.&lt;/li&gt;
&lt;li&gt;How words relate to each other.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It then uses these learned patterns to predict and generate language.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 How Does a Language Model Learn?
&lt;/h3&gt;

&lt;p&gt;Give thousands of examples. such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The cat is sleeping.&lt;/li&gt;
&lt;li&gt;The dog is running.&lt;/li&gt;
&lt;li&gt;I am drinking water.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Over time, the following patterns are learnt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cat and dog are animals.&lt;/li&gt;
&lt;li&gt;Sleeping and running are actions.&lt;/li&gt;
&lt;li&gt;Words follow certain grammatical patterns.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A language model learns language patterns from text in a similar broad sense, although its internal learning process is based on neural networks and mathematical optimization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example of next-word prediction:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prompt&lt;/th&gt;
&lt;th&gt;Predicted next word&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The sky is&lt;/td&gt;
&lt;td&gt;blue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;I am going to the&lt;/td&gt;
&lt;td&gt;market&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Please open the&lt;/td&gt;
&lt;td&gt;door&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  1.3 What Does "Predict the Likelihood of Next Set of Words" Mean?
&lt;/h3&gt;

&lt;p&gt;A language model looks at the words already written and calculates how likely different next words or tokens are.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Input: &lt;code&gt;I am going to the&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The model might assign probabilities like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Possible next word&lt;/th&gt;
&lt;th&gt;Illustrative likelihood&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;market&lt;/td&gt;
&lt;td&gt;35%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;office&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;park&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;school&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The model selects or samples a token according to its generation process, then continues predicting the next token.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.4 Language Model vs. Generative AI
&lt;/h3&gt;

&lt;p&gt;This is an important distinction in your notes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Language model&lt;/strong&gt; - A model that learns patterns in language and predicts language.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generative AI&lt;/strong&gt; - AI that can create new content, such as text, images, music, and videos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM (Large Language Model)&lt;/strong&gt; - A language model trained on a large amount of data, with many learned parameters, to understand and generate language.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  1.5 Modern LLMs
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Are Text prediction machines&lt;/li&gt;
&lt;li&gt;Recursive completion – generate text based on previously generated text hence can do this infinitely&lt;/li&gt;
&lt;li&gt;Generates text based on parameters – temperatures&lt;/li&gt;
&lt;li&gt;Uses training data for prediction&lt;/li&gt;
&lt;li&gt;LLM cannot do anything other except generate text&lt;/li&gt;
&lt;li&gt;However, tools can be used to perform external operations using the generated output&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Two Types of Language Modeling Tasks
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 Autoencoding
&lt;/h3&gt;

&lt;p&gt;A neural network architecture that learns to understand and represent data by encoding it into essential features and reconstructing the original input. In language modeling, it learns by predicting a missing word in a sentence using the surrounding context. Example: "The cat sat on the ___." → "mat".&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 Autoregressive
&lt;/h3&gt;

&lt;p&gt;A language modeling approach that predicts the next word or token based on the previous words. It generates text one token at a time. Example: "The cat sat on the" → "mat". ChatGPT and other GPT-style LLMs use autoregressive language modeling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Remember:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Autoencoding: Predict the missing word → Understand language.&lt;/li&gt;
&lt;li&gt;Autoregressive: Predict the next word → Generate language.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Autoregression can also be used for stock price and weather prediction, where future values are predicted from past values.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLM Training Methodology
&lt;/h2&gt;

&lt;p&gt;3 main stages that enable an LLM to perform useful tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.1 Learn/Model Language
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;First learn/model language&lt;/li&gt;
&lt;li&gt;To model language, the LLM is trained on very large amounts of text so it can learn patterns, relationships, syntax, context, and knowledge contained in that data.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.2 Pre-training
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Initial phase of LLM training - where the model learns from a very large and diverse dataset containing billions or trillions of tokens.&lt;/li&gt;
&lt;li&gt;Develop broad language understanding, contextual patterns, reasoning capabilities, and knowledge represented in the training data.&lt;/li&gt;
&lt;li&gt;For an autoregressive LLM, the model learns to predict the probability of the next token given the previous tokens across a huge number of text sequences.&lt;/li&gt;
&lt;li&gt;Training data is divided into sequences/chunks of tokens; each sequence provides many next-token prediction examples.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The model learns by adjusting its internal parameters (weights) based on the difference between its predictions and the actual training tokens.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Weights
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;a. They are numerical values (parameters) inside the neural network.&lt;/li&gt;
&lt;li&gt;b. During training, these parameters are continuously adjusted to capture statistical relationships between tokens and their contexts.&lt;/li&gt;
&lt;li&gt;c. For example, the model learns that "cat sat on the" is statistically more likely to be followed by "mat" than "house" in certain contexts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is to optimize the model parameters so that the model performs well on predicting tokens across diverse language examples.&lt;/p&gt;

&lt;p&gt;Training uses backpropagation and gradient descent to calculate how the weights should be adjusted, repeated across many training iterations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Training data can include:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a. Wikipedia and other reference sources&lt;/li&gt;
&lt;li&gt;b. Books and curated text datasets&lt;/li&gt;
&lt;li&gt;c. Web/internet text and other licensed or publicly available data&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.4 Pre-training Alone Is Not Enough
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;A pretrained model primarily learns to predict/continue text; it is not automatically a helpful conversational assistant.&lt;/li&gt;
&lt;li&gt;Model should respond appropriately to user requests and have useful conversations.&lt;/li&gt;
&lt;li&gt;This requires additional training to improve instruction following, helpfulness, safety, and response behavior.&lt;/li&gt;
&lt;li&gt;Pre-training is therefore more than simple phone autocomplete, but the underlying next-token prediction objective is similar.&lt;/li&gt;
&lt;li&gt;Additional instruction tuning/post-training is used to teach the model how to respond to instructions and conversations.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3.5 Fine-tuning
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Instruction Tuning
&lt;/h4&gt;

&lt;p&gt;Also called instruction fine-tuning (IFT); it is a form of supervised fine-tuning rather than simply "transfer learning."&lt;/p&gt;

&lt;p&gt;The model is trained to follow natural-language instructions and produce appropriate responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use an instruction dataset:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a. Contains instruction/prompt → response examples.&lt;/li&gt;
&lt;li&gt;b. The pretrained model already contains broad language patterns and learned representations, but needs examples of how to follow instructions and format useful responses.&lt;/li&gt;
&lt;li&gt;c. These datasets can be manually created or curated, often containing high-quality question-and-answer or task-and-response examples.&lt;/li&gt;
&lt;li&gt;d. Reinforcement Learning from Human Feedback (RLHF) can be an additional post-training stage used to align responses with human preferences; it is not the same thing as instruction tuning.&lt;/li&gt;
&lt;li&gt;e. A reward model can score responses based on learned human preferences, and reinforcement learning can optimize the model toward higher-reward behavior.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Fine-tuning
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Further training of an already pretrained/instruction-tuned model for a specific task, domain, behavior, or use case.&lt;/li&gt;
&lt;li&gt;Focused dataset relevant to the target task—for example, customer-support conversations, legal text, or a specific classification task.&lt;/li&gt;
&lt;li&gt;Additional training is performed on the existing model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Process:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a. Prepare examples representative of the desired task and behavior.&lt;/li&gt;
&lt;li&gt;b. Provide input/output examples showing what response or behavior is desired.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fine-tuning APIs offered by some model providers allow developers to upload training examples that teach the model "given this type of input, produce this type of output."&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokens and Embeddings
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Search engines primarily relied on keyword matching—pages containing more relevant occurrences of the search terms could rank higher.&lt;/li&gt;
&lt;li&gt;However, we want to search based on the meaning and context of words, known as &lt;strong&gt;Semantic search&lt;/strong&gt;, rather than exact word matching.&lt;/li&gt;
&lt;li&gt;Example: "Spring framework" should understand that Spring refers to the Java software framework, not the spring season or a car spring.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.1 How LLMs Process Tokens
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;LLMs ultimately process numbers, so text must first be converted into numerical representations.&lt;/li&gt;
&lt;li&gt;Character codes such as ASCII/Unicode are not sufficient because they represent characters, not learned linguistic relationships or meaning.&lt;/li&gt;
&lt;li&gt;LLMs learn language patterns and relationships from token sequences; semantic understanding emerges from the model's learned representations.&lt;/li&gt;
&lt;li&gt;Therefore, text is broken into manageable units called tokens, which the model can process efficiently.&lt;/li&gt;
&lt;li&gt;A token is a unit of text, not strictly a "unit of meaning"; meaning is represented through the model's learned embeddings and internal representations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is done using tokenization:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The tokenizer breaks text into tokens.&lt;/li&gt;
&lt;li&gt;Each token is assigned a unique token ID from the model's vocabulary.&lt;/li&gt;
&lt;li&gt;The model receives token IDs and processes them to predict/output token IDs.&lt;/li&gt;
&lt;li&gt;One token is not necessarily one word; it can be a word, part of a word, punctuation, whitespace, or other text fragment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.2 Visualizing Tokenization
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tokenizer playground:&lt;/strong&gt; GPT-4o, o200k_base&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt:&lt;/strong&gt; &lt;code&gt;i am doing tokenization&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token breakdown:&lt;/strong&gt; &lt;code&gt;i am doing token ization&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token count:&lt;/strong&gt; 5 tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token IDs:&lt;/strong&gt; &lt;code&gt;72, 939, 5306, 6602, 2860&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.3 Why Isn't One Word Always One Token?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Token IDs are assigned to vocabulary entries so the model can efficiently process text as numbers.&lt;/li&gt;
&lt;li&gt;Words are useful linguistic units, but they are not always the most efficient units for an LLM.&lt;/li&gt;
&lt;li&gt;Giving every possible word its own token would require a very large vocabulary.&lt;/li&gt;
&lt;li&gt;Related words such as learn, learning, learned can share subword pieces rather than requiring completely separate tokens.&lt;/li&gt;
&lt;li&gt;Languages contain an enormous number of words, forms, names, and new words.&lt;/li&gt;
&lt;li&gt;Subword tokenization allows the model to handle related and previously unseen words more efficiently.&lt;/li&gt;
&lt;li&gt;Words can also have multiple meanings; context helps the model determine the intended meaning.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.4 The Solution
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Break text into commonly occurring subword/character sequences, rather than relying only on complete words.&lt;/li&gt;
&lt;li&gt;Tokenizers generally use frequency/statistical patterns in training data to build an efficient vocabulary.&lt;/li&gt;
&lt;li&gt;Punctuation such as a comma can be its own token or part of another token.&lt;/li&gt;
&lt;li&gt;Frequently occurring text sequences can receive their own tokens, making processing more efficient.&lt;/li&gt;
&lt;li&gt;Tokenization is not necessarily based on grammatical word boundaries.&lt;/li&gt;
&lt;li&gt;Different models use different tokenizers and vocabularies, so the same text can produce different tokens and token counts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Overall process:&lt;/strong&gt; Text → Tokenizer → Token IDs → LLM → learned embeddings/internal representations → predicted Token IDs → Tokenizer → Text. Embeddings are the numerical vectors the model uses to represent tokens in a learned continuous space; they are different from token IDs.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.5 Token Decoding
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The tokenizer processes the text from left to right and splits it into tokens based on the tokenizer's vocabulary and tokenization rules.&lt;/li&gt;
&lt;li&gt;It generally prefers token pieces that efficiently match the available vocabulary; it is not simply a greedy search for the longest English word.&lt;/li&gt;
&lt;li&gt;Example: tokenization may be split into token + ization because those pieces exist in the tokenizer's vocabulary.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.6 What Numbers Are Assigned to Tokens
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Each token is assigned a token ID from the tokenizer's vocabulary. The ID itself does not store the meaning of the token.&lt;/li&gt;
&lt;li&gt;The token ID is then converted into an embedding vector, which is a learned &lt;strong&gt;set of numerical values [,,,]&lt;/strong&gt; helps the model represent relationships between tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4.7 How the Model Learns
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Example: Picture classification: how does AI determine whether a picture is a cat?&lt;/li&gt;
&lt;li&gt;The model learns to identify cat vs. not-cat from many examples.&lt;/li&gt;
&lt;li&gt;It adjusts its internal parameters based on more examples and learns which visual patterns are useful.&lt;/li&gt;
&lt;li&gt;Features: whiskers, ears, eyes, fur, and body shape can provide signals that a picture may contain a cat.&lt;/li&gt;
&lt;li&gt;Different features can have different importance; a feature such as whiskers may provide a stronger signal than something like background color.&lt;/li&gt;
&lt;li&gt;The model learns patterns by representing features numerically and combining many signals rather than relying on a single feature.&lt;/li&gt;
&lt;li&gt;The model examines features/patterns across many examples and learns which visual characteristics are associated with cats.&lt;/li&gt;
&lt;li&gt;It also learns from incorrect predictions, adjusting its internal parameters when its prediction differs from the correct answer.&lt;/li&gt;
&lt;li&gt;Over many examples, it learns that some attributes have a stronger statistical relationship with the target than others.&lt;/li&gt;
&lt;li&gt;For example, whiskers, ears, eyes, and body shape can be useful signals for identifying a cat, while features such as day vs. night or background are generally less reliable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Feature Matrix
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does AI require matrix multiplication?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In AI, matrix operations are fundamental because models need to process and transform large numbers of values efficiently.&lt;/li&gt;
&lt;li&gt;A feature can be represented as a numerical dimension in a vector; for example, whiskers can be one feature.&lt;/li&gt;
&lt;li&gt;For a simple example: whiskers = 0 for a spoon and whiskers = 1 for a cat or dog.&lt;/li&gt;
&lt;li&gt;Attributes/features can therefore be represented using numerical values for each object.&lt;/li&gt;
&lt;li&gt;In real AI models, the model learns useful features automatically rather than us manually defining features such as whiskers.&lt;/li&gt;
&lt;li&gt;During training, the model adjusts its parameters to capture patterns and correlations that help it perform its task.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5.1 Feature Matrix for Language
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;LLM create feature matrix for every token&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An LLM represents each token using a vector of learned numerical features.&lt;/li&gt;
&lt;li&gt;Columns represent learned dimensions/features, and rows represent tokens.&lt;/li&gt;
&lt;li&gt;Each token therefore has a value for every learned dimension.&lt;/li&gt;
&lt;li&gt;The value in each cell represents how strongly that token is represented along that particular learned dimension; these are embedding values, not simply weights.&lt;/li&gt;
&lt;li&gt;For example, a token such as "hello" may have a strong value along a learned dimension associated with greetings or conversational context.&lt;/li&gt;
&lt;li&gt;Because there can be hundreds or thousands of dimensions, each token is represented by an N-dimensional vector.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Embedding
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;[1, 1.04, -2.55, ...]&lt;/code&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Away of representing things using numbers/vectors based on learned patterns and relationships.&lt;/li&gt;
&lt;li&gt;This set of numerical values is called an embedding.&lt;/li&gt;
&lt;li&gt;Every token is mapped to an embedding vector that captures learned patterns and relationships.&lt;/li&gt;
&lt;li&gt;These embeddings are part of how the model represents and processes language; meaning is not stored in a single number or dimension.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  6.1 Understanding Embeddings Using Dimensions
&lt;/h3&gt;

&lt;p&gt;Imagine, for simplicity, that we have three dimensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dimension 1: Pet ↔ Not pet&lt;/li&gt;
&lt;li&gt;Dimension 2: Wild ↔ Not wild&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Dimension 3: Mammal ↔ Not mammal&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Each word/token receives a numerical value along each dimension.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;For example, dog, cat, lion, and crocodile would have different values across these dimensions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Dog and cat would tend to have more similar representations because they share several characteristics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;In real LLMs, there are many more dimensions—often hundreds or thousands—not just three manually defined dimensions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tokens with similar learned representations tend to be closer together in the embedding space.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tokens with different representations tend to be farther apart.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Therefore, similar concepts or usage patterns can occupy nearby regions of the embedding space.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  6.2 Embedding Math
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;Slide: Embedding math!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Embedding math!&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;v("king") − v("man") + v("woman") ≈ v("queen")&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;v("king") represents the embedding vector for king.&lt;/li&gt;
&lt;li&gt;The relationship captured between king/man and woman/queen can approximately appear through vector arithmetic.&lt;/li&gt;
&lt;li&gt;This demonstrates that embeddings can capture relationships and patterns between words/tokens.&lt;/li&gt;
&lt;li&gt;Embeddings influence how an LLM processes language and are an important part of how it develops useful representations of meaning and context.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Tokens and Their Values
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Each token has a number called a token ID (an identifier used by the model).&lt;/li&gt;
&lt;li&gt;The token ID is used to look up the token's embedding vector.&lt;/li&gt;
&lt;li&gt;An embedding is a set of numerical values representing how strongly a token is represented across the model's learned dimensions/features.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Q&amp;amp;A: tokenizer vocab splitting&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question 1&lt;/strong&gt;&lt;br&gt;
If our tokenizer's vocab contains the subwords {"uni", "vers", "ity", "un", "iverse"}, how would it most likely split the word "university"?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Answer&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://dev.togreedy%20longest-match"&gt;"uni", "vers", "ity"&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  7.1 Text Similarity
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;From tokens to text&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Problem statement&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Given a sequence of tokens x₁, x₂, …, xₜ, predict the most likely next token xₜ₊₁.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Two things need to happen:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The next token depends on the context of the preceding text, potentially including many earlier tokens.&lt;/li&gt;
&lt;li&gt;Not every part of the previous text is equally relevant to predicting the next token.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How do you know what's important?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The model learns during training which parts of the preceding context are relevant to other parts.&lt;/li&gt;
&lt;li&gt;During processing, mechanisms such as attention allow the model to give more importance to relevant tokens and less importance to less-relevant tokens.&lt;/li&gt;
&lt;li&gt;In this way, the model learns from patterns in the training data which information is useful for predicting the next token.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Transformer Architecture
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What transformers do&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;They transform text to an "intermediate" language (encoding)&lt;/li&gt;
&lt;li&gt;Language is embeddings + positional embeddings&lt;/li&gt;
&lt;li&gt;Get the best next embeddings&lt;/li&gt;
&lt;li&gt;Transform that back to text (decoding)&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;A Transformer is a neural-network architecture designed to solve the language-modeling problem of understanding context and predicting the next token.&lt;/li&gt;
&lt;li&gt;It uses attention to determine which other tokens are important to each token when building its contextual representation.&lt;/li&gt;
&lt;li&gt;This idea was introduced in the landmark paper "Attention Is All You Need."&lt;/li&gt;
&lt;li&gt;They transform tokens into context-aware numerical representations inside the network.&lt;/li&gt;
&lt;li&gt;The input representation combines token embeddings + positional information.&lt;/li&gt;
&lt;li&gt;Attention and other Transformer layers transform these representations to produce information useful for predicting the next token.&lt;/li&gt;
&lt;li&gt;The model converts the final numerical output into probabilities over possible tokens, and a token is selected/generated from those probabilities.&lt;/li&gt;
&lt;li&gt;Each token effectively "asks" other tokens how relevant they are to its current context through the attention mechanism.&lt;/li&gt;
&lt;li&gt;This is especially important because many words have multiple meanings, and their meaning depends on the surrounding context.&lt;/li&gt;
&lt;li&gt;The token's representation is updated through the Transformer layers based on its relationship with other tokens in the sentence/context.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Get the next best embedding&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Each token "asks" every other token "how relevant are you to me?"&lt;/li&gt;
&lt;li&gt;Builds a context-aware representation&lt;/li&gt;
&lt;li&gt;Goes through a process to decide what to do with the relevant tokens&lt;/li&gt;
&lt;li&gt;Goes through weight matrix to generate the final output tokens&lt;/li&gt;
&lt;li&gt;There are lot of ambiguous words in english language and meaning depends on the context&lt;/li&gt;
&lt;li&gt;Embedding is transformed on the fly or adjusted on the fly depending on the meaning of each token with rest pect to the other token in the sentence&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;"How relevant are you?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Consider these sentences&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The animal didn't cross the street because it was too tired&lt;/li&gt;
&lt;li&gt;The animal didn't cross the street because it was too wide&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;In the first sentence, "it" is likely referring to the animal, because an animal can be tired.&lt;/li&gt;
&lt;li&gt;In the second sentence, "it" is likely referring to the street, because a street can be wide.&lt;/li&gt;
&lt;li&gt;Attention helps the Transformer use these contextual relationships to build different representations for the same token depending on the surrounding text.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Key idea:&lt;/strong&gt; The embedding starts as a learned representation of a token, but Transformer layers transform it into a context-dependent representation based on the surrounding tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context Length in LLMs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;LLM have context limit&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLMs have a context limit — the maximum number of tokens the model can consider in a single request.&lt;/li&gt;
&lt;li&gt;The Transformer processes relationships between tokens using attention, and practical/model-design limits determine how many tokens can be handled at once.&lt;/li&gt;
&lt;li&gt;The context typically includes both input tokens and the tokens generated as output, subject to the model's total context window.&lt;/li&gt;
&lt;li&gt;Context length is therefore measured in tokens, not words or characters.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  9.1 LLM State
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;LLMs are stateless and do not have memory&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLMs are generally stateless between separate API requests; the model itself does not automatically remember previous conversations.&lt;/li&gt;
&lt;li&gt;Each new request is processed based on the information provided in that request and any external state/memory system.&lt;/li&gt;
&lt;li&gt;The trained model primarily consists of learned parameters (weights); conversation history is not stored inside those weights.&lt;/li&gt;
&lt;li&gt;Therefore, the model does not inherently retain your previous conversation after the request ends.&lt;/li&gt;
&lt;li&gt;To make an LLM behave as though it remembers a conversation, the application can send relevant previous conversation/context with each request.&lt;/li&gt;
&lt;li&gt;Conversation context can contain sensitive or private information, so applications need appropriate data-handling and privacy controls.&lt;/li&gt;
&lt;li&gt;As context becomes very large, cost, latency, and sometimes the model's ability to use all the information effectively can become concerns.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is RAG
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Retrieval Augmented Generation&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;A common way to handle knowledge that is too large, private, or frequently changing to put entirely into the LLM context.&lt;/li&gt;
&lt;li&gt;Dynamically retrieves relevant document chunks at query time.&lt;/li&gt;
&lt;li&gt;Retrieved text should be relevant and useful; the LLM uses the retrieved context to generate the answer, but retrieval quality still matters.&lt;/li&gt;
&lt;li&gt;Often used when the information changes frequently or comes from private documents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Goal of RAG&lt;/strong&gt; – retrieve relevant information that fits within the model's available context window and provide it to the LLM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; "Hello bot, what is the return policy for this particular product of company x?"&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The LLM may not have the company's latest or specific return policy in its model knowledge, so the application needs to retrieve the relevant information from the company's documents or knowledge base.&lt;/li&gt;
&lt;li&gt;The basic process is: find the relevant part of the policy → insert it into the LLM's context → provide the user's question → ask the LLM to generate an answer based on that retrieved context.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  10.1 Vectors and Chunks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The key challenge is: how do we find the specific portion of the documents that is relevant to the user's question?&lt;/li&gt;
&lt;li&gt;First, documents are typically split into smaller chunks that can be individually retrieved.&lt;/li&gt;
&lt;li&gt;When the user asks a question, the system searches for the chunks most relevant to that question.&lt;/li&gt;
&lt;li&gt;The relevant chunks are then inserted into the LLM's context along with the user's query.&lt;/li&gt;
&lt;li&gt;To find the appropriate chunks, we can use &lt;strong&gt;semantic similarity—representing&lt;/strong&gt; the query and &lt;strong&gt;document chunks as embeddings&lt;/strong&gt; and comparing their &lt;strong&gt;vectors&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The system &lt;strong&gt;generates an embedding&lt;/strong&gt; for the user's query and compares it with the embeddings of the document chunks.&lt;/li&gt;
&lt;li&gt;It retrieves the top-K most similar/relevant chunks and provides them to the LLM.&lt;/li&gt;
&lt;li&gt;This overall approach is called Retrieval-Augmented Generation (RAG):

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval:&lt;/strong&gt; find relevant information from the knowledge base.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Augmentation:&lt;/strong&gt; add that retrieved information to the user's query/context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation:&lt;/strong&gt; the LLM uses the augmented context to generate the response.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  RAG Pipeline
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;RAG Pipeline&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;User prompt&lt;/strong&gt; → User prompt arrives — this contains what the user is asking LLM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM&lt;/strong&gt; - The LLM or application creates a retrieval query, which may be the original prompt or a reformulated query to retrieve data from vector database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector DB&lt;/strong&gt; — The query is embedded and searched against the vector database. Top-k relevant chunks are retrieved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context assembly&lt;/strong&gt; — Retrieved chunks are added to the original prompt/context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Final generation&lt;/strong&gt; — Feed the assembled context to the LLM to generate the response.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Vector Databases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is a vector database?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A database optimized for storing and searching high-dimensional vectors, often alongside metadata or references to the original documents.&lt;/li&gt;
&lt;li&gt;Can retrieve records using IDs/metadata, but its key RAG capability is vector similarity search.&lt;/li&gt;
&lt;li&gt;Finds vectors that are nearest or most similar to a given query vector.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The problem&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Finding the right text is not always easy.&lt;/li&gt;
&lt;li&gt;Simple keyword/text matching may miss relevant content with different wording.&lt;/li&gt;
&lt;li&gt;Semantic/vector search can find text based on meaning and contextual similarity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The solution&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chunk the documents into smaller pieces.&lt;/li&gt;
&lt;li&gt;Generate an embedding for each document chunk.&lt;/li&gt;
&lt;li&gt;Generate an embedding for the user's query.&lt;/li&gt;
&lt;li&gt;Find the chunks whose vectors are most similar/relevant to the query vector.&lt;/li&gt;
&lt;li&gt;Add the retrieved chunks to the LLM's context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Vector databases help in these.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A vector database query&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"Which stored embedding vectors are closest to my query vector, and what documents do they point at?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  12.1 How It Works — Ingestion
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Precompute embeddings for each document/chunk using a chosen embedding model.&lt;/li&gt;
&lt;li&gt;Store the vector along with a unique ID and metadata/payload, such as document ID, source, or text.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  12.2 How It Works — Indexing
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The database builds a vector-search index, often using an Approximate Nearest Neighbor (ANN) algorithm.&lt;/li&gt;
&lt;li&gt;ANN structures organize/search vectors efficiently so similar vectors can be found much faster than comparing the query with every stored vector.&lt;/li&gt;
&lt;li&gt;Performance depends on the index and implementation; O(log N) is not guaranteed and should not be treated as a general property of vector databases.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  12.3 How It Works — Querying
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Generate an embedding for the incoming question/query.&lt;/li&gt;
&lt;li&gt;The vector index quickly returns the top-k most similar vectors, usually approximately rather than by exhaustive comparison.&lt;/li&gt;
&lt;li&gt;Retrieve their IDs/metadata and associated document chunks for use as context.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Code example: vector DB usage&lt;/p&gt;


&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;your_vector_db&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Client&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;your_embedding_model&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;embed&lt;/span&gt;

&lt;span class="c1"&gt;# 1. Initialize
&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Upsert documents
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;embed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upsert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Query
&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is the return policy for electronic items?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;q_vec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;embed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;q_vec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  12.4 Some Vector DBs
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Pinecone&lt;/li&gt;
&lt;li&gt;Qdrant&lt;/li&gt;
&lt;li&gt;Weaviate&lt;/li&gt;
&lt;li&gt;Milvus&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>Claude Certified Developer - Foundations certification Overview</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Sun, 13 Sep 2026 20:02:28 +0000</pubDate>
      <link>https://dev.to/yashnigam/claude-certified-developer-foundations-certification-overview-4n07</link>
      <guid>https://dev.to/yashnigam/claude-certified-developer-foundations-certification-overview-4n07</guid>
      <description>&lt;p&gt;I recently completed the Claude Certified Developer - Foundations certification.&lt;/p&gt;

&lt;p&gt;This certification is based on the official course from Anthropic: &lt;a href="https://anthropic-partners.skilljar.com/path/claude-certified-developer-foundations" rel="noopener noreferrer"&gt;https://anthropic-partners.skilljar.com/path/claude-certified-developer-foundations&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;However, this course is only available to Anthropic partners. Below is my overview of the modules and structure of the Prep Course.&lt;/p&gt;

&lt;h1&gt;
  
  
  Claude Certified Developer - Foundations Prep Course
&lt;/h1&gt;

&lt;p&gt;The course consists of the following main modules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;MSO Foundations&lt;/strong&gt;&lt;br&gt;
Learn the model fundamentals and technical foundations the rest of the Developer Foundations course builds on.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1.1 How LLMs Behave&lt;/li&gt;
&lt;li&gt;1.2 Models &amp;amp; Reasoning&lt;/li&gt;
&lt;li&gt;1.3 Prompting Modes&lt;/li&gt;
&lt;li&gt;1.4 Technical Substrate&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Production-Grade Prompting, Agents &amp;amp; Tool Use&lt;/strong&gt;&lt;br&gt;
Build your first production integration on Claude, with reliable prompts, tools, context management, and agent loops.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2.1 Prompting Craft&lt;/li&gt;
&lt;li&gt;2.2 Extended Thinking&lt;/li&gt;
&lt;li&gt;2.3 Tool-Use and Schema Design&lt;/li&gt;
&lt;li&gt;2.4 Streaming Responses&lt;/li&gt;
&lt;li&gt;2.5 Context Engineering&lt;/li&gt;
&lt;li&gt;2.6 Agent Construction&lt;/li&gt;
&lt;li&gt;2.7 Agent Memory&lt;/li&gt;
&lt;li&gt;2.8 Cumulative Debug Task&lt;/li&gt;
&lt;li&gt;2.9 Multimodal and Batch Ingestion&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Claude Code, MCP &amp;amp; Integration&lt;/strong&gt;&lt;br&gt;
Learn to make a working Claude integration configurable, shareable, and safe to connect to real systems.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;3.1 Module Orientation &amp;amp; Learning Objectives&lt;/li&gt;
&lt;li&gt;3.2 The Explore → Plan → Code Loop&lt;/li&gt;
&lt;li&gt;3.3 Permission Modes&lt;/li&gt;
&lt;li&gt;3.4 Settings Configuration Hierarchy&lt;/li&gt;
&lt;li&gt;3.5 Placing the Human Review Gate&lt;/li&gt;
&lt;li&gt;3.6 Durable Project Context: CLAUDE.md&lt;/li&gt;
&lt;li&gt;3.7 Rules Instruction Files&lt;/li&gt;
&lt;li&gt;3.8 Hooks&lt;/li&gt;
&lt;li&gt;3.9 Subagents&lt;/li&gt;
&lt;li&gt;3.10 Packaging Workflows: Skills&lt;/li&gt;
&lt;li&gt;3.11 Custom Commands &amp;amp; Plugin Packaging&lt;/li&gt;
&lt;li&gt;3.12 Building an MCP Server&lt;/li&gt;
&lt;li&gt;3.13 MCP Transport&lt;/li&gt;
&lt;li&gt;3.14 Prompt Caching (Applied to MCP/Tool Definitions)&lt;/li&gt;
&lt;li&gt;3.15 Retrieval-Augmented Generation (RAG) — Two Forms&lt;/li&gt;
&lt;li&gt;3.16 MCP Configuration Scope&lt;/li&gt;
&lt;li&gt;3.17 MCP Tool-Level Permissions&lt;/li&gt;
&lt;li&gt;3.18 GitHub MCP Server — Worked Authentication Example&lt;/li&gt;
&lt;li&gt;3.19 Enterprise Authentication Patterns&lt;/li&gt;
&lt;li&gt;3.20 Regulated/Enterprise Deployment Requirements&lt;/li&gt;
&lt;li&gt;3.21 Code Modernization as an Applied Case Study&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Production Engineering, Evals &amp;amp; Security&lt;/strong&gt;&lt;br&gt;
Learn to take an agent that works in development and prove it holds up under real production traffic.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;4.1 Module Introduction&lt;/li&gt;
&lt;li&gt;4.2 Evals &amp;amp; Judges&lt;/li&gt;
&lt;li&gt;4.3 Testing &amp;amp; Tracing&lt;/li&gt;
&lt;li&gt;4.4 Failure Handling &amp;amp; Model Selection&lt;/li&gt;
&lt;li&gt;4.5 Cost &amp;amp; Orchestration&lt;/li&gt;
&lt;li&gt;4.6 Security&lt;/li&gt;
&lt;li&gt;4.7 Cumulative Task&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Accelerators &amp;amp; IP Contribution&lt;/strong&gt;&lt;br&gt;
Package a build that works into one that survives reuse, review, and deployment beyond the engagement that created it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;5.1 Module Introduction&lt;/li&gt;
&lt;li&gt;5.2 Packaging for Reuse&lt;/li&gt;
&lt;li&gt;5.3 Contributing Back&lt;/li&gt;
&lt;li&gt;5.4 Requirements &amp;amp; Lifecycle&lt;/li&gt;
&lt;li&gt;5.5 Systems Lifecycle for Claude Applications&lt;/li&gt;
&lt;li&gt;5.6 Deployment &amp;amp; Versioning&lt;/li&gt;
&lt;li&gt;5.7 Comparing Platforms&lt;/li&gt;
&lt;li&gt;5.8 Trust Boundaries&lt;/li&gt;
&lt;li&gt;5.9 Cumulative Task&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Module 1: MSO Foundations
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1.1 How LLMs Behave
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1.1.1 Tokens&lt;/strong&gt;&lt;br&gt;
Everything Claude processes (prompt, history, tools, results, output) is measured in tokens, not words/characters, and both pricing and context budget are token-based. Output tokens cost more than input tokens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.1.2 Context Window&lt;/strong&gt;&lt;br&gt;
The max tokens allowed in a single request — system prompt, user prompt, documents, history, tool results, and output combined. Exceeding it returns a &lt;code&gt;model_context_window_exceeded&lt;/code&gt; stop reason.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.1.3 Sampling &amp;amp; Temperature&lt;/strong&gt;&lt;br&gt;
Claude samples the next token from a probability distribution rather than picking one fixed "best" token. Temperature (a request parameter) tunes this: lower = more repeatable, higher = more varied/creative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.1.4 Non-Determinism&lt;/strong&gt;&lt;br&gt;
Sampling means identical inputs can yield different, equally valid outputs, so exact-string-match tests are unreliable. Test for properties instead — e.g., "required field present" or "value within range."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.1.5 Testing Model Output: Structural vs. Semantic Correctness&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check Type&lt;/th&gt;
&lt;th&gt;Definition&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Structural correctness&lt;/td&gt;
&lt;td&gt;Deterministic, yes/no checks&lt;/td&gt;
&lt;td&gt;Regex match, valid JSON, exact substring, value within tolerance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic correctness&lt;/td&gt;
&lt;td&gt;Meaning-based checks, can't be scripted deterministically&lt;/td&gt;
&lt;td&gt;Summary captures key points, correct tone, accurate despite different phrasing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Semantic checks need an &lt;strong&gt;LLM-as-judge&lt;/strong&gt;: a separate model call scores the output, often against a rubric/reference answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.1.6 Evals&lt;/strong&gt;&lt;br&gt;
A testing framework for non-deterministic outputs: scores quality across many test cases and reports an aggregate pass rate (e.g., "87% passed"). Consists of test inputs, correctness criteria (structural or LLM-judge), and an aggregation method.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 Models &amp;amp; Reasoning
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1.2.1 Claude Model Family&lt;/strong&gt;&lt;br&gt;
Four tiers trading off cost, latency, capability, and quality:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet&lt;/td&gt;
&lt;td&gt;Balanced default for most production workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku&lt;/td&gt;
&lt;td&gt;Optimized for speed/cost within its capability range&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus&lt;/td&gt;
&lt;td&gt;For demanding work beyond Sonnet's envelope&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fable&lt;/td&gt;
&lt;td&gt;Highest-capability tier, for the hardest reasoning/coding/agentic tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1.2.2 Model Selection Strategy&lt;/strong&gt;&lt;br&gt;
Start with Sonnet by default; move up a tier only when an eval shows it fails the quality bar, or down to Haiku only when an eval shows the quality drop is acceptable. Model choice should be eval-driven, not assumed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.2.3 Reasoning Modes&lt;/strong&gt;&lt;br&gt;
Reasoning mode (on/off, separate from model choice) lets the model spend extra tokens "thinking" before answering; current models use adaptive thinking tuned via an effort setting (the older &lt;code&gt;budget_tokens&lt;/code&gt; control is deprecated, now returns a 400 error). Thinking is hidden by default and worth the cost only on hard, multi-step problems — not simple lookups.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.3 Prompting Modes
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1.3.1 Zero-shot / One-shot / Multi-shot (Few-shot) Prompting&lt;/strong&gt;&lt;br&gt;
Distinguished by how many worked examples are given in the prompt:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Examples Given&lt;/th&gt;
&lt;th&gt;Best Used When&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Zero-shot&lt;/td&gt;
&lt;td&gt;None (instruction only)&lt;/td&gt;
&lt;td&gt;Task is simple, output shape is obvious&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-shot&lt;/td&gt;
&lt;td&gt;One input/output example&lt;/td&gt;
&lt;td&gt;A single reference clarifies expected output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-shot (Few-shot)&lt;/td&gt;
&lt;td&gt;Several examples&lt;/td&gt;
&lt;td&gt;Needs specific structure, casing, or edge-case handling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1.3.2 Cost Tradeoff of Examples&lt;/strong&gt;&lt;br&gt;
Each example consumes tokens on every call and eats into context budget, so examples aren't free — they trade quality/precision against cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.3.3 Interaction with Model Choice&lt;/strong&gt;&lt;br&gt;
More capable models often succeed zero-shot where smaller ones need few-shot examples, so added examples can substitute for a cheaper model. Best practice: start with the simplest model and fewest examples meeting your eval bar, then add either only if needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.4 Technical Substrate
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1.4.1 SDK vs. Raw REST API&lt;/strong&gt;&lt;br&gt;
Claude is accessed via an HTTP REST API (JSON over your API key); the SDK just wraps auth, request construction, retries, and parsing. Both hit the same API and models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.4.2 Response Delivery Patterns&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Synchronous&lt;/td&gt;
&lt;td&gt;Send request, wait for full response, then act — simplest pattern&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Streaming&lt;/td&gt;
&lt;td&gt;Delivered in pieces via server-sent events as generated; output appears immediately, client reassembles the final message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Asynchronous (&lt;code&gt;AsyncAnthropic&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Non-blocking async/await enables concurrency without blocking, while each call still returns in real time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Message Batches API&lt;/td&gt;
&lt;td&gt;High-volume/offline: submit a batch, poll for completion; up to 24h latency for lower per-token cost&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Module 2: Production-Grade Prompting, Agents &amp;amp; Tool Use
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 Prompting Craft
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.1.1 Diagnosing Prompt Failures Instead of Adding Words&lt;/strong&gt;&lt;br&gt;
When a prompt fails, fix it by identifying which structural technique is missing — not by rewording or adding instructions. Rewording doesn't fix boundary confusion or format drift; only the matching technique does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.2 The Four Failure → Fix Mapping&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Missing Technique&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wrong output shape (prose instead of JSON/label)&lt;/td&gt;
&lt;td&gt;Output constraint&lt;/td&gt;
&lt;td&gt;Nothing specified the response's form/stopping point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content/scope drift over turns&lt;/td&gt;
&lt;td&gt;System prompt (or a more specific one)&lt;/td&gt;
&lt;td&gt;Behavioral contract too vague to hold across turns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Right task, invented structure&lt;/td&gt;
&lt;td&gt;Few-shot examples&lt;/td&gt;
&lt;td&gt;Claude can't infer exact structure from description alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Works on tested inputs, breaks on edge case&lt;/td&gt;
&lt;td&gt;Constraint covering that variant&lt;/td&gt;
&lt;td&gt;Prompt only validated against a narrow input set&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.1.3 System Prompts&lt;/strong&gt;&lt;br&gt;
Carry the persistent behavioral contract for the whole session — role, output format, rules that must not change between turns. Write once as a stable instruction layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.4 XML Tags&lt;/strong&gt;&lt;br&gt;
Used to separate instructions from examples/data (e.g., &lt;code&gt;&amp;lt;sample_input&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;ideal_output&amp;gt;&lt;/code&gt;) so Claude doesn't misread examples as part of the task itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.5 Few-Shot Examples&lt;/strong&gt;&lt;br&gt;
Show the exact input→output pattern (structure, casing, format) rather than describing it — closes gaps written instructions leave open, especially for edge cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.6 Output Constraints&lt;/strong&gt;&lt;br&gt;
Explicitly define the exact form of the response (e.g., "return only one label, no other text") — controls form independent of content, preventing parser-breaking variability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.7 When to Stack vs. Simplify vs. Diagnose&lt;/strong&gt;&lt;br&gt;
Stack all four techniques for complex tasks with defined output contracts and edge cases; simplify for simple tasks (e.g., plain summarization). If a prompt has grown longer over 3+ re-prompts without improving, stop and diagnose the missing technique instead of adding more text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.8 Worked Postmortem: Six-Pass Classification Prompt&lt;/strong&gt;&lt;br&gt;
A ticket-classifier prompt was revised 6 times, growing longer without fixing inconsistency — Pass 4 added description instead of a constraint, Pass 5's verbosity induced verbose output (latency regression, no accuracy gain). Pass 6 fixed it by replacing prose with an output constraint + 2 few-shot examples.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.1.9 Structured Outputs: Moving Control from Prompt to API&lt;/strong&gt;&lt;br&gt;
Instead of asking for a shape via prompt text, you give the API a JSON schema and the model is constrained at generation time (constrained decoding) — only schema-matching tokens can be produced.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;What It Does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;JSON outputs&lt;/td&gt;
&lt;td&gt;Set &lt;code&gt;output_config.format&lt;/code&gt; to &lt;code&gt;type: json_schema&lt;/code&gt; + your schema — constrains the final response text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strict tool use&lt;/td&gt;
&lt;td&gt;Set &lt;code&gt;strict: true&lt;/code&gt; on a tool definition — constrains/validates arguments before your code runs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Tradeoffs: first request on a new schema is slower (grammar compiled, cached 24h); input tokens rise slightly (format-description system prompt injected); a guaranteed schema doesn't guarantee success — check &lt;code&gt;stop_reason&lt;/code&gt; for &lt;code&gt;refusal&lt;/code&gt;/&lt;code&gt;max_tokens&lt;/code&gt;; incompatible with message prefilling.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 Extended Thinking
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.2.1 What It Does&lt;/strong&gt;&lt;br&gt;
When enabled, Claude reasons step-by-step in a separate thinking block before the final answer. On newest models this content is hidden by default — request a summarized display to see it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.2.2 Adaptive Thinking / Effort Setting&lt;/strong&gt;&lt;br&gt;
Enabled via the &lt;code&gt;thinking&lt;/code&gt; parameter (off by default); depth is tuned via an effort setting, not a fixed token budget. The older &lt;code&gt;budget_tokens&lt;/code&gt; control is deprecated (400 error on newest models), and thinking tokens bill the same as output tokens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.2.3 When to Use Extended Thinking&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Multi-step reasoning (math, multi-hop logic, dependent action planning)&lt;/td&gt;
&lt;td&gt;Enable, match effort to problem depth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mechanical tasks (classification, format conversion, lookups)&lt;/td&gt;
&lt;td&gt;Leave off — no benefit, wastes tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agentic loops planning across tool calls&lt;/td&gt;
&lt;td&gt;Enable, budget for the planning step&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.2.4 The Carry-Back Rule (Critical Constraint)&lt;/strong&gt;&lt;br&gt;
In tool-use loops with thinking enabled, every thinking block must be sent back unchanged next turn — each has a signature, and editing/dropping it (even redacted blocks) breaks the signature and the API rejects the request. Manage context bloat via context engineering, never by stripping thinking blocks.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 Tool-Use and Schema Design
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.3.1 Tool-Use and Schema Design&lt;/strong&gt;&lt;br&gt;
Covered indirectly via structured outputs/strict tool use (see 2.1.9); explicit schema-design content beyond tool descriptions affecting routing appears under Agent Construction (see 2.6).&lt;/p&gt;

&lt;h3&gt;
  
  
  2.4 Streaming Responses
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.4.1 Why Streaming&lt;/strong&gt;&lt;br&gt;
Sends the response in pieces as generated (via server-sent events) instead of waiting for the full message, removing blank-screen wait on long responses. Your code must reassemble blocks itself and handle early termination.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.4.2 Event Sequence and Handler Actions&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Handler Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;message_start&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;New message beginning&lt;/td&gt;
&lt;td&gt;Initialize empty content array&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;content_block_start&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;New block opening (text/tool_use/thinking)&lt;/td&gt;
&lt;td&gt;Create slot at that index&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;content_block_delta&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Incremental fragment of a block&lt;/td&gt;
&lt;td&gt;Append to block; tool_use JSON isn't parseable until block closes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;content_block_stop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Block complete&lt;/td&gt;
&lt;td&gt;Finalize block (first point tool_use JSON is parseable)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;message_delta&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Top-level changes (stop_reason, usage)&lt;/td&gt;
&lt;td&gt;Record stop_reason&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;message_stop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stream complete&lt;/td&gt;
&lt;td&gt;Assembled content is now the finished message&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.4.3 Never Act on a Partial Block&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;tool_use&lt;/code&gt; inputs arrive as fragmented JSON strings, invalid until &lt;code&gt;content_block_stop&lt;/code&gt;. Parsing/running tools before block closure causes malformed JSON or missing arguments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.4.4 Commit to History Only After &lt;code&gt;message_stop&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
Only add a streamed assistant turn to history once &lt;code&gt;message_stop&lt;/code&gt; arrives and every block is fully assembled — a turn from an interrupted stream has an incomplete &lt;code&gt;tool_use&lt;/code&gt; block and violates tool-use pairing rules on the next request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.4.5 Handling Interrupted Streams&lt;/strong&gt;&lt;br&gt;
Treat accumulated content as provisional until &lt;code&gt;message_stop&lt;/code&gt;; on interruption, discard the partial turn and retry. Check &lt;code&gt;stop_reason&lt;/code&gt; from &lt;code&gt;message_delta&lt;/code&gt; before continuing a loop — &lt;code&gt;tool_use&lt;/code&gt; means calls are ready to run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.4.6 Postmortem: "Read Loop Ended" ≠ "Message Complete"&lt;/strong&gt;&lt;br&gt;
A handler that commits history whenever its read loop exits (instead of gating on &lt;code&gt;message_stop&lt;/code&gt;) can silently commit corrupted &lt;code&gt;tool_use&lt;/code&gt; blocks — the validation error surfaces on the &lt;em&gt;next&lt;/em&gt; request, not the one that caused it.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.5 Context Engineering
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.5.1 Model Selection Recap&lt;/strong&gt;&lt;br&gt;
Four-tier family (Fable, Opus, Sonnet, Haiku) trading cost/latency/capability — start with Sonnet, move tiers only on eval results. Confirm current lineup/identifiers against docs at build time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.2 Context Window Is a Shared, Finite Budget&lt;/strong&gt;&lt;br&gt;
Covers system prompt, history, every persisting tool result, and output. An over-budget request is rejected pre-generation; hitting the ceiling mid-response returns partial output with &lt;code&gt;model_context_window_exceeded&lt;/code&gt; — neither path auto-trims, so the app must manage it, and production sessions fill context faster than dev/test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.3 Four Strategies for Managing Context Budget&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;What It Does&lt;/th&gt;
&lt;th&gt;When to Use&lt;/th&gt;
&lt;th&gt;What's Lost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pruning&lt;/td&gt;
&lt;td&gt;Rewind to an earlier message, drop everything after&lt;/td&gt;
&lt;td&gt;After an unproductive/dead-end path&lt;/td&gt;
&lt;td&gt;Everything after the rewind point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compaction (&lt;code&gt;/compact&lt;/code&gt;; server-side beta or manual summarization)&lt;/td&gt;
&lt;td&gt;Summarizes history into a condensed version&lt;/td&gt;
&lt;td&gt;Approaching context ceiling, want to retain knowledge&lt;/td&gt;
&lt;td&gt;Any detail not captured in the summary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clearing (&lt;code&gt;/clear&lt;/code&gt;; new API session)&lt;/td&gt;
&lt;td&gt;Starts fresh, empty context&lt;/td&gt;
&lt;td&gt;Next task is unrelated&lt;/td&gt;
&lt;td&gt;All session context (persist elsewhere, e.g. CLAUDE.md)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subagent handoffs&lt;/td&gt;
&lt;td&gt;Spawn isolated subagent with task-specific context; returns summary&lt;/td&gt;
&lt;td&gt;Self-contained delegable subtasks&lt;/td&gt;
&lt;td&gt;Visibility into subagent's intermediate reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.5.4 Prompt Caching &amp;amp; Token Counting&lt;/strong&gt;&lt;br&gt;
Prompt caching reuses processing of a stable prefix at a fraction of cost via &lt;code&gt;cache_control&lt;/code&gt; (&lt;code&gt;type: ephemeral&lt;/code&gt;, up to 4 breakpoints) — the highest-leverage cost reduction for multi-turn sessions. &lt;code&gt;count_tokens&lt;/code&gt; returns a request's token count without running inference, for dev validation and production budget gating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.5 RAG: Three Failure Points&lt;/strong&gt;&lt;br&gt;
RAG lets an LLM retrieve external documents via chunking → embedding match → assembly → answer. Chunking must balance size (too small loses context, too large dilutes matches); embedding match can miss exact-term matches (pair with lexical search); sloppy assembly causes the model to ignore retrieved content entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.6 Indexed vs. Iterative RAG&lt;/strong&gt;&lt;br&gt;
Indexed (fetch-once) is inspectable/testable but needs an index to build/maintain/secure — good for a stable corpus. Iterative (search-across-rounds) avoids staleness but costs more tokens/time and is less inspectable — better for changing corpora or multi-step questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.7 Compaction: Summarizer Prompt Design&lt;/strong&gt;&lt;br&gt;
A vague summarizer prompt ("summarize so far") loses task-critical detail; a well-specified one (preserve file paths, decisions, errors/resolutions) retains what matters. Under-specified summarizers are a leading cause of multi-session agent failures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.8 Subagent Handoffs for Long-Horizon Tasks&lt;/strong&gt;&lt;br&gt;
Instead of growing the context window, decompose the task and give each subagent only a scoped task, minimal context, relevant prior results, needed tools, and clear exit conditions. Keeps per-turn cost low at the expense of implementation overhead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.5.9 Postmortem: Context Budget Not Tested Against Production Data&lt;/strong&gt;&lt;br&gt;
A receipt-processing agent's 40k-token self-imposed cap worked fine against 800-token/call dev fixtures but was exhausted by turn 8 against real 3,200-token production tool outputs, crowding out system instructions. Fix: prune tool outputs after use and compact proactively — always measure actual production-scale token costs before shipping.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.6 Agent Construction
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.6.1 Definition&lt;/strong&gt;&lt;br&gt;
An agent is a multi-step tool-use loop with managed context and a defined goal. Key upfront question: does the problem actually require an agent, given the added coordination overhead, context cost, and failure surface?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.6.2 Workflow vs. Agent Decision&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Choose a Workflow When...&lt;/th&gt;
&lt;th&gt;Choose an Agent When...&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Steps can be enumerated in code&lt;/td&gt;
&lt;td&gt;Path can't be enumerated in advance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error cost is high, step-level guardrails needed&lt;/td&gt;
&lt;td&gt;Non-determinism acceptable, actions constrained by toolset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard observability tooling required&lt;/td&gt;
&lt;td&gt;Inputs vary unpredictably&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inputs are well-constrained&lt;/td&gt;
&lt;td&gt;Task requires creative tool sequencing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.6.3 Three Wiring Paths&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Who Runs the Loop&lt;/th&gt;
&lt;th&gt;You Own&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Raw API loop&lt;/td&gt;
&lt;td&gt;Your code&lt;/td&gt;
&lt;td&gt;Everything: loop, execution, context mgmt, retries, exit conditions&lt;/td&gt;
&lt;td&gt;Full control / learning / library constraints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent SDK&lt;/td&gt;
&lt;td&gt;SDK, in your process&lt;/td&gt;
&lt;td&gt;Tool execution + app; SDK gives loop structure, context mgmt, tool registration&lt;/td&gt;
&lt;td&gt;Claude Code's scaffolding without rebuilding it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Managed Agents (public beta)&lt;/td&gt;
&lt;td&gt;Anthropic (server-side, via SSE)&lt;/td&gt;
&lt;td&gt;App layer + agent definition as a versioned resource&lt;/td&gt;
&lt;td&gt;Long-running tasks (mins–hours), avoid building sandbox/loop&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.6.4 Managed Agents Specifics&lt;/strong&gt;&lt;br&gt;
Anthropic runs the loop/sandbox/retries server-side with stateful stored sessions; not currently eligible for Zero Data Retention or HIPAA BAA (rules out PHI/ZDR workloads). Requires the &lt;code&gt;managed-agents-2026-04-01&lt;/code&gt; beta header; common path is prototyping on Agent SDK then re-expressing config as a Managed Agents resource for production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.6.5 Four Steps Common to Every Agent Loop&lt;/strong&gt;&lt;br&gt;
(1) Register tools with consistent schema, (2) set a system prompt scoped to the agent's task/toolset, (3) handle the tool-use loop — every &lt;code&gt;tool_use&lt;/code&gt; block must get a &lt;code&gt;tool_result&lt;/code&gt; and all blocks from one turn resolved together, (4) define explicit exit conditions rather than relying on Claude to volunteer completion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.6.6 Human-in-the-Loop (HITL) Insertion Points&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Insertion Point&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;th&gt;Risk Addressed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before destructive tool call&lt;/td&gt;
&lt;td&gt;Write/delete/send operation&lt;/td&gt;
&lt;td&gt;High — irreversible actions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After a planning step&lt;/td&gt;
&lt;td&gt;Plan generated, about to execute&lt;/td&gt;
&lt;td&gt;Medium — wrong plan even if execution is correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On unexpected output&lt;/td&gt;
&lt;td&gt;Error flag, empty result, out-of-bounds value&lt;/td&gt;
&lt;td&gt;Variable — catches failures retries won't fix&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.6.7 Tool Orchestration: Over-Tooling vs. Under-Tooling&lt;/strong&gt;&lt;br&gt;
Too many overlapping tools causes erratic routing (the more common production problem, from registering tools "just in case"); too few causes hallucinated paths. Start with the minimum tool set and add only on confirmed capability gaps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.6.8 Regulated Data Constraints Determine Delivery Route&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;th&gt;Rules Out&lt;/th&gt;
&lt;th&gt;Survives Review&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Attorney-client privilege&lt;/td&gt;
&lt;td&gt;Unaudited consumer Claude.ai calls&lt;/td&gt;
&lt;td&gt;Direct API/SDK via firm's SSO-authenticated app, routed through firm-approved LLM gateway with full logging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HIPAA (PHI)&lt;/td&gt;
&lt;td&gt;Any endpoint without a covering BAA&lt;/td&gt;
&lt;td&gt;BAA-covered direct API, or Bedrock/Vertex on a HIPAA-eligible account (BAA excludes Console, Workbench, betas, consumer plans)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPR / data residency&lt;/td&gt;
&lt;td&gt;Routes without pinned execution region; direct API (no EU residency)&lt;/td&gt;
&lt;td&gt;Bedrock/Vertex with region pinned to the covered jurisdiction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FedRAMP / government&lt;/td&gt;
&lt;td&gt;Any non-authorized endpoint, incl. dev/test on commercial endpoint&lt;/td&gt;
&lt;td&gt;Claude for Government (C4G), Bedrock GovCloud, Vertex Assured Workloads (Claude Enterprise on AWS Marketplace is NOT FedRAMP authorized)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internal data-residency policy&lt;/td&gt;
&lt;td&gt;Any vendor outside the approved list&lt;/td&gt;
&lt;td&gt;Delivery route on the approved vendor only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(SOC 2 governs system operation, not endpoint selection — covered in Module 4.)&lt;/p&gt;

&lt;h3&gt;
  
  
  2.7 Agent Memory
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.7.1 Memory Scope&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Persists&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Use When&lt;/th&gt;
&lt;th&gt;Lost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;In-context&lt;/td&gt;
&lt;td&gt;Within a single session&lt;/td&gt;
&lt;td&gt;Zero retrieval overhead, token cost grows with conversation&lt;/td&gt;
&lt;td&gt;Short sessions fitting fully in context&lt;/td&gt;
&lt;td&gt;Everything at session end&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;External storage&lt;/td&gt;
&lt;td&gt;Across sessions/users/instances, in a DB&lt;/td&gt;
&lt;td&gt;Retrieval latency + read/write engineering&lt;/td&gt;
&lt;td&gt;Cross-session continuity needed&lt;/td&gt;
&lt;td&gt;Nothing (cost is latency/complexity)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Summarized memory&lt;/td&gt;
&lt;td&gt;Condensed version injected next session&lt;/td&gt;
&lt;td&gt;Lower cost than full replay, drops detail&lt;/td&gt;
&lt;td&gt;Long-running conversations exceeding budget&lt;/td&gt;
&lt;td&gt;Anything summarizer didn't preserve&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stateless (none)&lt;/td&gt;
&lt;td&gt;Nothing&lt;/td&gt;
&lt;td&gt;Zero overhead&lt;/td&gt;
&lt;td&gt;Self-contained, one-off jobs&lt;/td&gt;
&lt;td&gt;All prior context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.7.2 Design-Time Decision, Not Refactor-Time&lt;/strong&gt;&lt;br&gt;
Choosing memory scope during design avoids costly production refactors; a common failure pattern defaults to full in-context history, which works until token cost/latency climbs and hits the context ceiling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.7.3 Skills: Reusable, On-Demand Instruction Sets&lt;/strong&gt;&lt;br&gt;
A Skill is a reusable &lt;code&gt;SKILL.md&lt;/code&gt; (frontmatter: name, description + instructions body) that Claude loads only when the description matches an incoming request — unlike in-context memory, instructions aren't resident every session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.7.4 Skills vs. CLAUDE.md vs. In-Context Instructions&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Loads&lt;/th&gt;
&lt;th&gt;Context Cost&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Skill&lt;/td&gt;
&lt;td&gt;On-demand, when description matches&lt;/td&gt;
&lt;td&gt;Low (only name+description loaded upfront)&lt;/td&gt;
&lt;td&gt;Task-specific expertise not needed every session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CLAUDE.md&lt;/td&gt;
&lt;td&gt;Every session unconditionally (CLI); controlled by &lt;code&gt;settingSources&lt;/code&gt; in Agent SDK&lt;/td&gt;
&lt;td&gt;Fixed overhead per session&lt;/td&gt;
&lt;td&gt;Always-on project-wide standards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In-context instructions&lt;/td&gt;
&lt;td&gt;Every turn in that session&lt;/td&gt;
&lt;td&gt;Grows with session length, doesn't survive session end&lt;/td&gt;
&lt;td&gt;Short, one-off sessions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.7.5 Skills on the Messages API (Beta)&lt;/strong&gt;&lt;br&gt;
Requires two beta headers (&lt;code&gt;code-execution-2025-08-25&lt;/code&gt;, &lt;code&gt;skills-2025-10-02&lt;/code&gt;); Skills run inside the code execution container, not the app's own environment, affecting available tools/filesystem access.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.7.6 Subagents and Skills&lt;/strong&gt;&lt;br&gt;
Subagents don't automatically inherit Skills or conversation history from the parent (clean context on delegation), but do inherit the parent's permission context. A subagent needing a Skill must have it explicitly listed in its own configuration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.7.7 Postmortem: In-Context Memory Filling by Session Four&lt;/strong&gt;&lt;br&gt;
An escalation-support agent worked fine in dev (long continuous sessions) but failed in production, where many shorter sessions accumulated state until injected history exceeded 40k tokens by session 4 — before any tool call. Fix: refactor to external storage, though this took longer under production pressure than doing it at design time.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.8 Cumulative Debug Task
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.8.1 Cumulative Debug Task&lt;/strong&gt;&lt;br&gt;
Applied exercise combining prior concepts, not new theory. Observed bug-diagnosis mapping: vague tool descriptions → schema/routing failure; interrupted streams/stripped thinking blocks → carry-back rule violation; orphaned &lt;code&gt;tool_result&lt;/code&gt; without matching &lt;code&gt;tool_use&lt;/code&gt; → context/pairing violation; concatenated full session transcripts → memory scope failure at scale.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.9 Multimodal and Batch Ingestion
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2.9.1 Image Token Cost&lt;/strong&gt;&lt;br&gt;
Images process in 28×28-pixel patches: cost = ⌈width/28⌉ × ⌈height/28⌉ visual tokens (e.g., 1000×1000px ≈ 1,296 tokens). Each tier has a max native resolution (oversized images downscaled first) — confirm current limits against docs, and measure production image token cost before building a pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.9.2 Three Ways to Send an Image/File&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Overhead&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inline base64&lt;/td&gt;
&lt;td&gt;Encode bytes directly in the message&lt;/td&gt;
&lt;td&gt;Full payload sent every request&lt;/td&gt;
&lt;td&gt;One-off images unlikely to be reused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;URL reference&lt;/td&gt;
&lt;td&gt;Pass a public URL; Claude fetches it&lt;/td&gt;
&lt;td&gt;No payload, but URL must stay stable/public/reachable&lt;/td&gt;
&lt;td&gt;Already-hosted, stable public images&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Files API (beta; not on Bedrock/Vertex)&lt;/td&gt;
&lt;td&gt;Upload once → get &lt;code&gt;file_id&lt;/code&gt; → reference later&lt;/td&gt;
&lt;td&gt;One-time upload cost, near-zero overhead thereafter&lt;/td&gt;
&lt;td&gt;Reused assets, large files, multi-turn conversations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.9.3 Sending PDFs&lt;/strong&gt;&lt;br&gt;
Uses a document block type (vs. image) with the same source structure (base64/URL/file_id); no required &lt;code&gt;name&lt;/code&gt; field, optional &lt;code&gt;title&lt;/code&gt;/&lt;code&gt;context&lt;/code&gt; fields, same token-cost and Files API reuse mechanics.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.9.4 Prompting Technique Carryover to Multimodal&lt;/strong&gt;&lt;br&gt;
The same four prompting techniques apply to image/PDF analysis, but images introduce visual ambiguity (overlapping objects, depth, occlusion) that prompts should explicitly instruct how to handle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.9.5 Message Batches API&lt;/strong&gt;&lt;br&gt;
For high-volume async processing — up to 100,000 requests or 256MB per batch. Submit → get &lt;code&gt;batch_id&lt;/code&gt; → poll → download results (arbitrary order, matched via &lt;code&gt;custom_id&lt;/code&gt;). Lower per-token cost, latency up to 24h — suited to offline pipelines/evals/bulk jobs, not real-time interactions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.9.6 Postmortem: "Chunked Loop" Mistaken for Batching&lt;/strong&gt;&lt;br&gt;
Looping over a list calling the synchronous API per item (even in smaller chunks) is not batching — it's serialized calls still hitting the same rate limits at volume. True batching requires the actual Batch API.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2.9.7 Use-Case Fit&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;API&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;User uploads photo, expects immediate result&lt;/td&gt;
&lt;td&gt;Synchronous&lt;/td&gt;
&lt;td&gt;Real-time required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nightly job classifying 5,000 records&lt;/td&gt;
&lt;td&gt;Batches API&lt;/td&gt;
&lt;td&gt;No latency constraint; cost savings matter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval run against 2,000 examples&lt;/td&gt;
&lt;td&gt;Batches API&lt;/td&gt;
&lt;td&gt;Offline, no real-time need&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chatbot reply generation&lt;/td&gt;
&lt;td&gt;Synchronous&lt;/td&gt;
&lt;td&gt;User is actively waiting&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;2.9.8 Two Failure Modes Combining Multimodal + Batch&lt;/strong&gt;&lt;br&gt;
(1) Misreading latency needs — using batch for a user-facing image flow fails because the user is waiting; (2) underestimating context cost — multiple large images/PDFs per request can blow past token limits at scale.&lt;/p&gt;




&lt;h2&gt;
  
  
  Module 3: Claude Code, MCP &amp;amp; Integration
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 Module Orientation &amp;amp; Learning Objectives
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.1.1 Module Orientation&lt;/strong&gt;&lt;br&gt;
Claude Code is a terminal-native dev partner running the same agent loop as the API, plus a permission layer, config system, and team-sharing features; MCP enables secure external-service integration. Core theme: configs that "work on your machine" often fail once shared, deployed, or run against production.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 The Explore → Plan → Code Loop
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.2.1 The Three-Phase Loop&lt;/strong&gt;&lt;br&gt;
Claude Code works in three phases: Explore (reads files, traces logic, no edits), Plan (proposes a structured edit description, requires human approval), Code (writes/executes changes only after approval) — this sequence reduces wrong assumptions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.2.2 Plan Mode as the Hook Point&lt;/strong&gt;&lt;br&gt;
Plan mode holds the agent in the explore phase, blocking all edits/commands until released — a good default for unfamiliar or high-stakes codebases.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Permission Modes
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.3.1 Permission Modes&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Auto-Approves&lt;/th&gt;
&lt;th&gt;Still Gated&lt;/th&gt;
&lt;th&gt;Limitations&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Default&lt;/td&gt;
&lt;td&gt;Reads only&lt;/td&gt;
&lt;td&gt;All edits/commands&lt;/td&gt;
&lt;td&gt;Safe but slow; baseline for new/unfamiliar projects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AcceptEdits&lt;/td&gt;
&lt;td&gt;Reads, file edits, common filesystem commands (&lt;code&gt;mkdir&lt;/code&gt;, &lt;code&gt;touch&lt;/code&gt;, &lt;code&gt;rm&lt;/code&gt;, &lt;code&gt;rmdir&lt;/code&gt;, &lt;code&gt;mv&lt;/code&gt;, &lt;code&gt;cp&lt;/code&gt;, &lt;code&gt;sed&lt;/code&gt;) in working dir&lt;/td&gt;
&lt;td&gt;All other shell commands; writes outside working dir; protected paths&lt;/td&gt;
&lt;td&gt;Good for trusted local work; not for running scripts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan&lt;/td&gt;
&lt;td&gt;Reads only&lt;/td&gt;
&lt;td&gt;All edits/commands until plan approved&lt;/td&gt;
&lt;td&gt;Exploration on sensitive/unfamiliar code; not for tasks needing output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auto&lt;/td&gt;
&lt;td&gt;Everything, but a classifier reviews each action and blocks escalation/hostile intent&lt;/td&gt;
&lt;td&gt;Production deploys, migrations, mass deletes, credential exfiltration, force-push to main&lt;/td&gt;
&lt;td&gt;Research preview — not a safety guarantee; varies by plan/model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DontAsk&lt;/td&gt;
&lt;td&gt;Only pre-approved allow-listed tools + read-only commands&lt;/td&gt;
&lt;td&gt;Everything else auto-denied&lt;/td&gt;
&lt;td&gt;Built for locked-down CI/scripts, not for reducing local friction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BypassPermissions&lt;/td&gt;
&lt;td&gt;All tool calls, no prompts, no safety checks&lt;/td&gt;
&lt;td&gt;Nothing (except catastrophic &lt;code&gt;rm -rf /&lt;/code&gt; or &lt;code&gt;rm -rf ~&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Only safe in isolated/disposable containers — never on a live dev workstation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  3.4 Settings Configuration Hierarchy
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.4.1 Configuration Hierarchy&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Use For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;User&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.claude/settings.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Every project on the machine, not committed&lt;/td&gt;
&lt;td&gt;Personal defaults (e.g., preferred mode)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.claude/settings.json&lt;/code&gt; (committed)&lt;/td&gt;
&lt;td&gt;Everyone who clones the repo&lt;/td&gt;
&lt;td&gt;Team-wide conventions, allow/deny rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local project&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.claude/settings.local.json&lt;/code&gt; (git-ignored)&lt;/td&gt;
&lt;td&gt;Personal overrides for one project&lt;/td&gt;
&lt;td&gt;Individual preferences not meant for the team&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;managed-settings.json&lt;/code&gt; (admin-set)&lt;/td&gt;
&lt;td&gt;Cannot be overridden by users/projects&lt;/td&gt;
&lt;td&gt;Org-wide security controls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.4.2 Rule Precedence&lt;/strong&gt;&lt;br&gt;
A deny rule always wins over an allow rule regardless of mode; enterprise-level deny rules are the most durable — they can't be removed by any developer and apply even under bypass mode.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.5 Placing the Human Review Gate
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.5.1 The Governing Question &amp;amp; Gate Placement&lt;/strong&gt;&lt;br&gt;
Ask "what's the worst outcome if this action runs unchecked?" — let low-stakes reversible actions through, gate hard-to-undo/sensitive-path actions via deny rules or prompts, and always require human review before merging changes to team-flagged sensitive code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.5.2 Postmortem: BypassPermissions Removed a Safety Prompt&lt;/strong&gt;&lt;br&gt;
Bypass mode skipped the confirmation prompt and protected-path guard that would have caught an overly broad file-pattern match, causing accidental deletion of production config files. Mitigation: set explicit deny rules on sensitive directories first; prefer classifier-gated modes (e.g., Auto) over full bypass.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.6 Durable Project Context: CLAUDE.md
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.6.1 CLAUDE.md Basics &amp;amp; &lt;code&gt;/init&lt;/code&gt;&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;CLAUDE.md&lt;/code&gt; at the project root loads into every session, prepended before any user message, so conventions/constraints/commands persist without restating; &lt;code&gt;/init&lt;/code&gt; generates a starter file by scanning the codebase (validate before relying on it).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.6.2 Size Dilution Failure Mode&lt;/strong&gt;&lt;br&gt;
As the file grows, each instruction becomes a smaller fraction of loaded context, reducing the chance any single rule is followed — keep &lt;code&gt;CLAUDE.md&lt;/code&gt; to behavior-changing constraints and move everything else into on-demand Skills.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.6.3 Postmortem: 847-Line CLAUDE.md&lt;/strong&gt;&lt;br&gt;
A real path restriction (line 347) was diluted among hundreds of unrelated lines (historical logs, archived notes) and the agent violated it. Fix: keep it as a working rule set, move path-specific rules to rules files, historical context to reference docs, non-negotiable constraints to hooks.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.7 Rules Instruction Files
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.7.1 Rules Files: Path-Scoped Context&lt;/strong&gt;&lt;br&gt;
Live in &lt;code&gt;.claude/rules/&lt;/code&gt;, scoped via a &lt;code&gt;paths&lt;/code&gt; glob in YAML frontmatter, loading into context only when Claude works with matching files — avoiding &lt;code&gt;CLAUDE.md&lt;/code&gt;-style dilution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.7.2 Scoping Comes from Frontmatter, Not Location&lt;/strong&gt;&lt;br&gt;
A rules file without a &lt;code&gt;paths&lt;/code&gt; field loads unconditionally at launch (same priority as &lt;code&gt;CLAUDE.md&lt;/code&gt;) regardless of subdirectory; pattern is broad/universal → &lt;code&gt;CLAUDE.md&lt;/code&gt;, narrow/path-specific → rules files (e.g., &lt;code&gt;.claude/rules/database.md&lt;/code&gt; with &lt;code&gt;paths: ["src/db/**/*.sql"]&lt;/code&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  3.8 Hooks
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.8.1 Hooks: Deterministic Lifecycle Control&lt;/strong&gt;&lt;br&gt;
Hooks intercept/control tool calls at fixed lifecycle points deterministically, unlike a &lt;code&gt;CLAUDE.md&lt;/code&gt; instruction the model might follow inconsistently — defined in settings files, configured via &lt;code&gt;/hooks&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.8.2 Hook Events&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Timing&lt;/th&gt;
&lt;th&gt;Can Block?&lt;/th&gt;
&lt;th&gt;Use For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PreToolUse&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Before tool call executes&lt;/td&gt;
&lt;td&gt;Yes — exit code 2 blocks it, stderr shown to agent&lt;/td&gt;
&lt;td&gt;Access control enforcement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PostToolUse&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;After tool call completes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Formatting, tests, audit logging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;UserPromptSubmit&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;On prompt submission, before processing&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Inject context, validate request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Stop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;When model finishes responding&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Notifications, cleanup, audit commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Notification&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;On Claude Code notifications (permission requests, 60s idle)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Route to external channel/logging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SessionStart&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Session start/resume&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Initialize state, validate env vars&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SessionEnd&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Session end&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Teardown, final audit writes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.8.3 Hook vs. Convention&lt;/strong&gt;&lt;br&gt;
A &lt;code&gt;PreToolUse&lt;/code&gt; hook enforces a constraint (e.g., blocking edits to a production config path) at every tool call, every session, regardless of permission mode — the difference between a guardrail (hook) and a convention (&lt;code&gt;CLAUDE.md&lt;/code&gt; instruction).&lt;/p&gt;

&lt;h3&gt;
  
  
  3.9 Subagents
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.9.1 Subagents: Isolated Context&lt;/strong&gt;&lt;br&gt;
Specialized assistants that run tasks in an isolated context — no inheritance of main conversation history, accumulated files, or session state; return only their output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.9.2 Built-in vs. Custom Subagent Behavior&lt;/strong&gt;&lt;br&gt;
Built-in &lt;code&gt;Explore&lt;/code&gt;/&lt;code&gt;Plan&lt;/code&gt; subagents skip &lt;code&gt;CLAUDE.md&lt;/code&gt; and git status (optimized for speed, so project rules don't apply); &lt;code&gt;general-purpose&lt;/code&gt; loads both. Custom subagents don't auto-inherit skills — list needed skills explicitly in the subagent's frontmatter (&lt;code&gt;.claude/agents&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.9.3 Mechanism Comparison&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Loads&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;Context Cost&lt;/th&gt;
&lt;th&gt;Belongs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CLAUDE.md&lt;/td&gt;
&lt;td&gt;Full file, prepended&lt;/td&gt;
&lt;td&gt;Every session&lt;/td&gt;
&lt;td&gt;Persistent, dilutes with size&lt;/td&gt;
&lt;td&gt;Universal constraints/commands&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rules file&lt;/td&gt;
&lt;td&gt;File contents&lt;/td&gt;
&lt;td&gt;On matching file access (or session start if unscoped)&lt;/td&gt;
&lt;td&gt;Path-scoped: only when triggered; unscoped: same as CLAUDE.md&lt;/td&gt;
&lt;td&gt;Path-specific guidance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hook&lt;/td&gt;
&lt;td&gt;Runs a script&lt;/td&gt;
&lt;td&gt;At configured lifecycle event&lt;/td&gt;
&lt;td&gt;Minimal&lt;/td&gt;
&lt;td&gt;Guardrails, automation, audit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subagent&lt;/td&gt;
&lt;td&gt;Task context only&lt;/td&gt;
&lt;td&gt;When dispatched&lt;/td&gt;
&lt;td&gt;Returns summary only&lt;/td&gt;
&lt;td&gt;Exploration/investigation, parallelizable work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  3.10 Packaging Workflows: Skills
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.10.1 Skills: Portable Markdown Procedures&lt;/strong&gt;&lt;br&gt;
A Skill is a portable &lt;code&gt;SKILL.md&lt;/code&gt; in &lt;code&gt;.claude/skills&lt;/code&gt; (frontmatter identifies/describes it, body holds steps); the same file can run in Claude Code, via Messages API, or via Agent SDK — only where it runs and what it can access changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.10.2 Skill Runtimes&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runtime&lt;/th&gt;
&lt;th&gt;How It Loads&lt;/th&gt;
&lt;th&gt;Where Steps Run&lt;/th&gt;
&lt;th&gt;Key Requirement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Filesystem discovery (description match or invoke by name)&lt;/td&gt;
&lt;td&gt;Local terminal, under active permission mode/deny rules&lt;/td&gt;
&lt;td&gt;Filesystem-based, governed by settings layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Messages API&lt;/td&gt;
&lt;td&gt;Sent with request, runs in code execution container&lt;/td&gt;
&lt;td&gt;Anthropic's container, not local machine&lt;/td&gt;
&lt;td&gt;Requires code-execution + skills beta headers; no local file/tool assumptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent SDK&lt;/td&gt;
&lt;td&gt;Loaded by agent; controlled by &lt;code&gt;settingSources&lt;/code&gt; (TS) / &lt;code&gt;setting_sources&lt;/code&gt; (Python)&lt;/td&gt;
&lt;td&gt;The SDK's own process, once filesystem sources are enabled&lt;/td&gt;
&lt;td&gt;Must explicitly set &lt;code&gt;settingSources&lt;/code&gt; — no reliable default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Managed Agents&lt;/td&gt;
&lt;td&gt;Defined once as an API resource (model, prompt, tools, MCP, skills)&lt;/td&gt;
&lt;td&gt;Anthropic-provisioned sandbox&lt;/td&gt;
&lt;td&gt;Requires &lt;code&gt;managed-agents-2026-04-01&lt;/code&gt; beta header; not ZDR/HIPAA-eligible; skills attached at agent-definition time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.10.3 Three Portability Rules&lt;/strong&gt;&lt;br&gt;
(1) Write the skill's description as a precise matching criterion; (2) don't assume local filesystem/tools exist — document dependencies explicitly; (3) subagents don't inherit skills in any runtime — must be explicitly listed.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.11 Custom Commands &amp;amp; Plugin Packaging
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.11.1 Custom Commands (Legacy vs. Skills)&lt;/strong&gt;&lt;br&gt;
Skills are now the recommended format for both explicit (&lt;code&gt;/skill-name&lt;/code&gt;) and automatic invocation; the older &lt;code&gt;.claude/commands/&lt;/code&gt; format still works but is legacy. Use &lt;code&gt;disable-model-invocation: true&lt;/code&gt; in frontmatter for commands that should only run when explicitly called.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.11.2 Plugin Namespacing&lt;/strong&gt;&lt;br&gt;
A plugin name becomes the command prefix (e.g., &lt;code&gt;/payments:run-tests&lt;/code&gt;), preventing collisions across plugins; renaming a plugin renames all its commands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.11.3 Plugins &amp;amp; Marketplaces&lt;/strong&gt;&lt;br&gt;
A plugin is a versioned bundle of skills, hooks, subagents, and MCP servers distributed via a marketplace. Anthropic's official marketplace is available by default; third-party ones are added via &lt;code&gt;/plugin marketplace add &amp;lt;owner/repo&amp;gt;&lt;/code&gt;. Enterprise admins can deploy plugins org-wide via managed settings, gated by a managed marketplace allowlist, paired with &lt;code&gt;extraKnownMarketplaces&lt;/code&gt; to auto-register for all users; managed-scope settings sit above user/project settings and cannot be overridden.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.11.4 Packaging Decision Table&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Reach For When&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Skill&lt;/td&gt;
&lt;td&gt;Task-specific procedure should stay out of context until needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom command&lt;/td&gt;
&lt;td&gt;Predictable, explicit high-frequency entry point wanted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plugin&lt;/td&gt;
&lt;td&gt;A working local setup needs to be shared/versioned across a team&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.11.5 Postmortem: Plugin Portability Failure&lt;/strong&gt;&lt;br&gt;
A plugin skill referenced an author's absolute local path and an undocumented env var — installed fine everywhere, but execution failed for every teammate. Fix: use &lt;code&gt;$CLAUDE_PROJECT_DIR&lt;/code&gt; / &lt;code&gt;${CLAUDE_PLUGIN_ROOT}&lt;/code&gt; for paths, bundle dependent assets, document/validate env vars at install, and test on a clean machine; note a locally-relied-on deny rule/hook isn't auto-included in a plugin bundle.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.12 Building an MCP Server
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.12.1 What MCP Is&lt;/strong&gt;&lt;br&gt;
MCP (Model Context Protocol) separates tool definitions from individual applications into a standalone server process exposing tools, resources, and prompts to any connecting client — build once, reuse everywhere.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.12.2 Tools, Resources, Prompts&lt;/strong&gt;&lt;br&gt;
Tools are actions the model can call; Resources are read-only data fetched directly into context (not via a tool call) — direct (fixed address) or templated (parameterized), used when cheap predictable injection beats a tool call (client support varies); Prompts are pre-written, vetted instruction templates exposed by name for consistent wording across clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.13 MCP Transport
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.13.1 Transport Options&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Transport&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Use For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;td&gt;Local subprocess via standard input/output&lt;/td&gt;
&lt;td&gt;Local tools, personal scripts, dev servers on your own machine — not shareable/remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP&lt;/td&gt;
&lt;td&gt;Standard network connection via URL&lt;/td&gt;
&lt;td&gt;Recommended for any non-local server — shared team servers, hosted integrations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SSE&lt;/td&gt;
&lt;td&gt;Legacy transport predating HTTP&lt;/td&gt;
&lt;td&gt;No longer recommended for new servers; treat as legacy if encountered&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.13.2 Context Cost &amp;amp; Tool Discovery&lt;/strong&gt;&lt;br&gt;
Connected servers' tool definitions would occupy context if loaded upfront; Claude Code defers loading and searches to discover/load only relevant tools per task by default (an opt-in mode loads upfront if definitions fit within ~10% of the context window). Connect only needed servers to keep requests lean.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.14 Prompt Caching (Applied to MCP/Tool Definitions)
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.14.1 How Prompt Caching Works&lt;/strong&gt;&lt;br&gt;
Caches processing of a stable request prefix so follow-up requests reuse it at a fraction of cost; the first request writes to cache, and an exact match is required — a single changed character before the cache point invalidates it. Best candidates: long system prompts, large tool-definition sets, frequently-queried reference docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.14.2 Cache Configuration Details&lt;/strong&gt;&lt;br&gt;
Enabled via &lt;code&gt;cache_control: {type: ephemeral}&lt;/code&gt; on the last block to cache (up to 4 breakpoints); request order is fixed (tools → system prompt → messages), so a breakpoint after tools caches definitions while messages stay dynamic. Default lifetime is 5 minutes from last read; opt-in 1-hour via &lt;code&gt;ttl: "1h"&lt;/code&gt;; a minimum token threshold applies (1,024 for most current models).&lt;/p&gt;

&lt;h3&gt;
  
  
  3.15 Retrieval-Augmented Generation (RAG) — Two Forms
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.15.1 Classical RAG vs. Agentic Search&lt;/strong&gt;&lt;br&gt;
Classical RAG pre-chunks and embeds source material into a database, matched by similarity at query time (index built in advance); agentic search has no pre-built index — the model searches/fetches live sources on demand (e.g., MCP tool discovery, Projects surfacing relevant sections). Both retrieve a relevant slice and generate from it — the difference is timing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.15.2 Two Properties of Retrieval&lt;/strong&gt;&lt;br&gt;
(1) Scales flat — request cost stays roughly constant as source material grows, since only a relevant slice is retrieved per query; (2) quality depends on retrievability — if retrieval misses the needed document the model never sees it, so well-named, well-organized files matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.16 MCP Configuration Scope
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.16.1 Scope Levels&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Applies To&lt;/th&gt;
&lt;th&gt;Use For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;~/.claude.json&lt;/code&gt; (per-project entry)&lt;/td&gt;
&lt;td&gt;Current project only, not shared&lt;/td&gt;
&lt;td&gt;Project-specific, not-yet-shared config&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User&lt;/td&gt;
&lt;td&gt;Personal Claude settings&lt;/td&gt;
&lt;td&gt;All your projects, still personal&lt;/td&gt;
&lt;td&gt;Personal utilities used everywhere&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.mcp.json&lt;/code&gt; (committed to repo)&lt;/td&gt;
&lt;td&gt;Everyone who clones the repo&lt;/td&gt;
&lt;td&gt;Team-wide shared server access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Centrally managed config (admin-controlled)&lt;/td&gt;
&lt;td&gt;Entire organization&lt;/td&gt;
&lt;td&gt;Shared internal services, security tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.16.2 Note on Project-Scoped stdio Servers&lt;/strong&gt;&lt;br&gt;
A project-scoped stdio server still runs from each teammate's own machine — each clone needs the runtime (e.g., Node for an &lt;code&gt;npx&lt;/code&gt;-launched server) installed locally.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.17 MCP Tool-Level Permissions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.17.1 Per-Tool Permission Rules&lt;/strong&gt;&lt;br&gt;
MCP tools are identified as &lt;code&gt;mcp__server__tool&lt;/code&gt;, so rules can target individual tools rather than the whole server — an allow rule on one tool lets it run without prompting while others on the same server still prompt; a deny rule on a write-capable tool blocks it while read-only tools remain available, and deny always overrides allow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.17.2 API MCP Connector: Scope vs. Governance&lt;/strong&gt;&lt;br&gt;
An &lt;code&gt;mcp_toolset&lt;/code&gt; object with a per-tool &lt;code&gt;enabled&lt;/code&gt; flag controls whether the model even sees a tool (context/scope control), distinct from a permission rule controlling whether an exposed tool may run (governance control) — often used together; verify exact syntax/beta headers against docs.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.18 GitHub MCP Server — Worked Authentication Example
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.18.1 GitHub MCP Setup&lt;/strong&gt;&lt;br&gt;
Transport: HTTP (remote, hosted by GitHub), registered via server URL; scope: Project for whole-team access, local for individual use; auth: a Personal Access Token passed as a Bearer token, supplied via environment variable and referenced in config — never committed inline (a committed token enters repo history permanently).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.18.2 OAuth Alternative&lt;/strong&gt;&lt;br&gt;
For services authenticating individual users via browser sign-in (e.g., Linear MCP) — client redirects to the provider's sign-in page, token issued/stored automatically after approval, no manual credential handling; right pattern when authorization is tied to user identity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.18.3 MCP Setup Reference&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Transport&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Secrets Handling&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Personal local tool&lt;/td&gt;
&lt;td&gt;stdio&lt;/td&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;Env vars only, never in config file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared team server&lt;/td&gt;
&lt;td&gt;HTTP&lt;/td&gt;
&lt;td&gt;Project (&lt;code&gt;.mcp.json&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;OAuth or env vars; never committed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Personal experiment&lt;/td&gt;
&lt;td&gt;stdio/HTTP&lt;/td&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;Env vars only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Org-wide deployment&lt;/td&gt;
&lt;td&gt;HTTP&lt;/td&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Admin-managed secrets, config locked&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  3.19 Enterprise Authentication Patterns
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.19.1 Auth Method by Service Type&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service Type&lt;/th&gt;
&lt;th&gt;Auth Method&lt;/th&gt;
&lt;th&gt;Secret Location&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Remote, user identity (SaaS/cloud)&lt;/td&gt;
&lt;td&gt;OAuth (server returns 401 → browser sign-in)&lt;/td&gt;
&lt;td&gt;Token issued/stored by client&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remote, service identity (internal API)&lt;/td&gt;
&lt;td&gt;API key via environment variable&lt;/td&gt;
&lt;td&gt;Environment only, injected by CI/pipeline runner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local, filesystem access&lt;/td&gt;
&lt;td&gt;stdio, no network auth&lt;/td&gt;
&lt;td&gt;Filesystem permission model + deny rules as governance&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.19.2 Secret Management: Three Practices&lt;/strong&gt;&lt;br&gt;
(1) Separation — config holds only a variable reference, never the value, since committed values persist in repo history; (2) storage — env var for local/short-lived secrets, a secret store for shared/audited ones (enables single-point rotation); (3) rotation — replace on a schedule and immediately after suspected exposure, scoped narrowly per credential.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.20 Regulated/Enterprise Deployment Requirements
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.20.1 Beyond Authentication: Three Requirements&lt;/strong&gt;&lt;br&gt;
(1) Configuration lock — enterprise managed settings prevent developers from overriding auth setup; (2) audit logging — a &lt;code&gt;PostToolUse&lt;/code&gt; hook logging every tool call/parameters, firing deterministically and unskippable by the model; (3) data residency — an HTTP endpoint pinned to a region plus a platform deployment enforcing regional processing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.20.2 Authentication/Integration Checklist&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service Type&lt;/th&gt;
&lt;th&gt;Auth&lt;/th&gt;
&lt;th&gt;Secrets&lt;/th&gt;
&lt;th&gt;Logging&lt;/th&gt;
&lt;th&gt;Config Lock&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Remote, user identity&lt;/td&gt;
&lt;td&gt;OAuth&lt;/td&gt;
&lt;td&gt;Client-stored token&lt;/td&gt;
&lt;td&gt;PostToolUse hook&lt;/td&gt;
&lt;td&gt;Enterprise managed settings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Remote, service identity&lt;/td&gt;
&lt;td&gt;API key (env var)&lt;/td&gt;
&lt;td&gt;Environment only&lt;/td&gt;
&lt;td&gt;PostToolUse hook&lt;/td&gt;
&lt;td&gt;Enterprise managed settings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;Filesystem permissions&lt;/td&gt;
&lt;td&gt;None needed&lt;/td&gt;
&lt;td&gt;PostToolUse hook&lt;/td&gt;
&lt;td&gt;Deny rules in managed settings&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;3.20.3 Postmortem: OAuth Staging-to-Production Failure&lt;/strong&gt;&lt;br&gt;
OAuth redirect URIs are registered per host, so a working staging connection doesn't guarantee production works — the production host's URI isn't automatically authorized, and regulated customers often require separate app registrations per environment. Fix: add the new host's redirect URI before cutover and include this in the deployment checklist.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.21 Code Modernization as an Applied Case Study
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;3.21.1 Modernization Risk Profile&lt;/strong&gt;&lt;br&gt;
Legacy modernization concentrates high blast radius, unpredictable dependencies, and limited reversibility; the Explore/Plan/Code loop and Plan mode hold the agent in read-only review, hooks guard sensitive paths during high-risk phases, and CLAUDE.md carries target-pattern conventions so the agent doesn't drift back to legacy patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3.21.2 Three Scoping Questions&lt;/strong&gt;&lt;br&gt;
Before high-risk agentic work, ask: (1) what's the blast radius — which systems depend on this code; (2) how are changes audited — is a &lt;code&gt;PostToolUse&lt;/code&gt; hook logging every tool call sufficient; (3) who approves each phase before the next begins — Plan mode enforces the explore/execute boundary, but the approval process itself must be defined beforehand.&lt;/p&gt;




&lt;h2&gt;
  
  
  Module 4: Production Engineering, Evals &amp;amp; Security
&lt;/h2&gt;

&lt;h3&gt;
  
  
  4.1 Module Introduction
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.1.1 Core Theme: Production Reveals What Development Hides&lt;/strong&gt;&lt;br&gt;
Success criteria never written down, retriable errors never given a path, budgets never instrumented, action boundaries never enforced — these are decisions, not bugs, and must be made on paper before failures happen live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.1.2 The Design Document — Four Decisions&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;What It Defines&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Success criteria&lt;/td&gt;
&lt;td&gt;Concrete, checkable output definitions (e.g., "a two-sentence summary listing every action item and owner") — what the eval set is built from&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure handling&lt;/td&gt;
&lt;td&gt;Expected production errors, each marked retriable/terminal, and what the user sees on unrecoverable failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost and latency budget&lt;/td&gt;
&lt;td&gt;Per-request budget, monthly cost ceiling, latency target, minimum reliability floor — set before architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trust boundary&lt;/td&gt;
&lt;td&gt;Which inputs are untrusted, and the smallest set of actions/access the feature needs — turns least privilege into an enforceable hook, not a remembered setting&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  4.2 Evals &amp;amp; Judges
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.2.1 What an Eval Is&lt;/strong&gt;&lt;br&gt;
A fixed set of input cases + expected behavior + grading, run and averaged into a trackable score — turns "done" from a feeling into a number. Write the eval before building the feature, forcing success to be defined upfront.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.2 Minimal Eval Pipeline&lt;/strong&gt;&lt;br&gt;
Load dataset → run each case → grade output → average scores. Change one variable (prompt, tool, or model) at a time between runs so you know what caused any score movement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.3 Three Grading Methods&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Catches&lt;/th&gt;
&lt;th&gt;Blind Spot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exact/string match&lt;/td&gt;
&lt;td&gt;Output with exactly one correct form&lt;/td&gt;
&lt;td&gt;Wrong answer, zero ambiguity, near-zero cost&lt;/td&gt;
&lt;td&gt;Fails on any valid paraphrase or reordering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code-graded check&lt;/td&gt;
&lt;td&gt;Structured/code output&lt;/td&gt;
&lt;td&gt;Invalid JSON, unparseable code, out-of-range values, missing fields&lt;/td&gt;
&lt;td&gt;Confirms well-formedness only, not content quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM-as-judge&lt;/td&gt;
&lt;td&gt;Open-ended quality (faithfulness, instruction-following, tone)&lt;/td&gt;
&lt;td&gt;What no code rule can express&lt;/td&gt;
&lt;td&gt;Noisy, costly, produces a confident-looking but meaningless number until calibrated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;4.2.4 Cost Dimension&lt;/strong&gt;&lt;br&gt;
Exact match/code checks run locally at near-zero cost (thousands per change); a judge is a second model call per case — grade format/structure with code on every commit, reserve judge calls for slower scheduled quality passes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.5 Building/Calibrating a Judge&lt;/strong&gt;&lt;br&gt;
Prompt the judge for strengths, weaknesses, and reasoning alongside the score, or it drifts to a "safe middle" (~6) regardless of quality. Calibrate by running it against a human-labeled set and measuring agreement — low agreement means an untrustworthy score, fixed by tightening the rubric.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.6 Coverage Over Perfection&lt;/strong&gt;&lt;br&gt;
A larger, noisier automated eval set usually reveals more than a small hand-graded one; use Claude to generate additional edge cases from a small labeled starting set, then spot-check them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.7 Iteration Workflow&lt;/strong&gt;&lt;br&gt;
Set goal → write initial prompt → run eval → read per-case failures → apply one change → re-run. Per-case breakdown matters as much as the average (can hide fixes canceling breaks); categorize failure cause (formatting, retrieval, long-input) to target the next fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.2.8 Postmortem: Field-Extraction Two-Date Failure&lt;/strong&gt;&lt;br&gt;
A feature passed ~12 manual checks (all single-date messages) and shipped with shape-only validation; a message with two dates extracted the wrong one, since validation confirmed shape, not correctness, and no holdout set covered that case. Fix: define graded cases including model-generated edge cases before shipping — the eval guards against regressions, it doesn't fix the underlying prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.3 Testing &amp;amp; Tracing
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.3.1 Four Test Levels&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Isolates&lt;/th&gt;
&lt;th&gt;Cannot Catch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Unit&lt;/td&gt;
&lt;td&gt;One function (parser, tool wrapper) alone&lt;/td&gt;
&lt;td&gt;How components fit together&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Functional&lt;/td&gt;
&lt;td&gt;One Claude call returns expected shape for an input&lt;/td&gt;
&lt;td&gt;Failures in the surrounding system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integration&lt;/td&gt;
&lt;td&gt;The handoff/seam between two components (e.g., retrieval → model)&lt;/td&gt;
&lt;td&gt;Whole-flow behavior emerging only end-to-end&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;End-to-end&lt;/td&gt;
&lt;td&gt;Full flow as a user would run it&lt;/td&gt;
&lt;td&gt;Where the break is — slowest, hardest to localize&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;4.3.2 Integration Seam Failures&lt;/strong&gt;&lt;br&gt;
Most silent production failures live at the integration seam — each side passes its own tests while the handoff between them is broken (e.g., one returns a list of dicts, the other expects a plain string).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.3.3 Tracing&lt;/strong&gt;&lt;br&gt;
Records each run step (prompt, tool calls, intermediate outputs, timing), turning "the case failed" into "step 4: parser raised KeyError on a field the model didn't return" — a 5-minute fix instead of a day of investigation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.3.4 Retrieval Routing (Cost-Aware)&lt;/strong&gt;&lt;br&gt;
A cheap classifier routes simple lookups to fetch-once retrieval and multi-part questions to iterative/agentic search, avoiding defaulting everything to the expensive or shallow path. Skip the router if all traffic is one shape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.3.5 Postmortem: Unit/Functional Passed, End-to-End Failed&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;retrieve()&lt;/code&gt; returned a list of dicts while &lt;code&gt;build_prompt()&lt;/code&gt; expected a plain string, causing malformed context and the model to answer from memory instead of retrieved content — only an integration test exercising the real handoff would catch this.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.4 Failure Handling &amp;amp; Model Selection
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.4.1 Retriable vs. Terminal&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;HTTP Status Codes&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retriable&lt;/td&gt;
&lt;td&gt;429, 529, 500, 502, 503, 504&lt;/td&gt;
&lt;td&gt;Rate limit, overload, transient server fault, timeout&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal&lt;/td&gt;
&lt;td&gt;400, 401, 403, 404&lt;/td&gt;
&lt;td&gt;Bad request, auth failure, missing resource, permissions problem&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Misclassifying as terminal fails loudly and gets fixed (safe default when unsure); misclassifying as retriable hammers the service and hides the real problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.4.2 SDK-Level Retries&lt;/strong&gt;&lt;br&gt;
Anthropic client libraries auto-retry transient failures with progressive delays up to a configurable cap — avoid stacking a custom retry loop on top (multiplies attempts against a rate limit); either let the SDK own retries or turn them off and own the whole path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.4.3 &lt;code&gt;retry-after&lt;/code&gt; Header&lt;/strong&gt;&lt;br&gt;
Present on 429/529 responses, gives the exact wait time — more precise than blind backoff. Read it first; fall back to exponential backoff only when absent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.4.4 Tool Errors &amp;amp; Refusals&lt;/strong&gt;&lt;br&gt;
Tool errors must return to Claude explicitly with &lt;code&gt;is_error: true&lt;/code&gt; — a silent empty result is treated as valid data, producing a confident-but-wrong answer. A refusal (&lt;code&gt;stop_reason: "refusal"&lt;/code&gt;) is a 200 at the HTTP layer, so the retriable classifier won't catch it — treat it as terminal, raise and log, never retry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.4.5 Error-Handling Decision Table&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Error&lt;/th&gt;
&lt;th&gt;Retriable?&lt;/th&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Fallback&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rate limit (429)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Exponential backoff + jitter, honor &lt;code&gt;retry-after&lt;/code&gt;, capped attempts&lt;/td&gt;
&lt;td&gt;Clean error or cached/simpler result after cap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overloaded (529)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Backoff (Anthropic-side load, not a rate-limit signal)&lt;/td&gt;
&lt;td&gt;Fail over or graceful error if persists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bad request (400)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No retry&lt;/td&gt;
&lt;td&gt;Fix/reject input, surface to caller&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool result error&lt;/td&gt;
&lt;td&gt;Depends&lt;/td&gt;
&lt;td&gt;Retry only if cause is transient&lt;/td&gt;
&lt;td&gt;Return error flag to Claude, never silence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refusal (200, &lt;code&gt;stop_reason: refusal&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No retry&lt;/td&gt;
&lt;td&gt;Raise to caller, log, never treat as valid output&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;4.4.6 Postmortem: Unhandled Rate Limit + Retry Storm&lt;/strong&gt;&lt;br&gt;
A loop with no error handling worked in low-volume dev; the first production 429 raised an unhandled exception, and the instinct to add immediate retries deepened the rate limit. Fix: honor &lt;code&gt;retry-after&lt;/code&gt; first, fall back to capped exponential backoff with jitter, and fail fast on terminal statuses (400/401/403/404).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.4.7 Model Selection in Production&lt;/strong&gt;&lt;br&gt;
Family order: Fable (hardest reasoning/coding/agentic) → Opus (above Sonnet's envelope) → Sonnet (balanced default) → Haiku (speed/cost-optimized). Start with Sonnet, move up only when an eval shows a missed quality bar, move down only when an eval shows an acceptable drop — defaulting to the most capable model "just in case" is the most common, expensive mistake. Route bulk traffic to a default model, override via a cheap signal (task type, length, difficulty) only where needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.5 Cost &amp;amp; Orchestration
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.5.1 Observability: Three Metrics&lt;/strong&gt;&lt;br&gt;
Instrument token usage (input/output), latency, and error rate per call from the start — per-call logging turns "why is the bill high?" into "which step, on which request type, is responsible?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.2 Cost/Latency Levers: Model Selection &amp;amp; Streaming&lt;/strong&gt;&lt;br&gt;
Route simpler work to smaller/faster models, reserve capable ones for steps that need them. With streaming + tool use, accumulate &lt;code&gt;content_block_delta&lt;/code&gt; events by index and never act on a &lt;code&gt;tool_use&lt;/code&gt; block until the stream closes — a broken mid-stream response needs a full retry, not partial output passed downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.3 Prompt Caching Economics&lt;/strong&gt;&lt;br&gt;
Cache writes cost a premium (1.25x base for 5-min TTL, 2x for 1-hour); cache reads cost ~0.1x standard input — only pays off when reads outnumber writes. Two modes: automatic (single flag, system manages breakpoints) or explicit &lt;code&gt;cache_control&lt;/code&gt; breakpoints.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.4 Message Batches API&lt;/strong&gt;&lt;br&gt;
Async processing at lower per-request cost in exchange for non-immediate completion — right for non-urgent, high-volume work (overnight jobs, backfills), wrong for anything a user is waiting on; compounds with prompt caching when a batch reuses shared context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.5 Multi-Agent Orchestration (Orchestrator-Worker)&lt;/strong&gt;&lt;br&gt;
A lead agent decomposes a task, delegates to parallel subagents, then synthesizes results — genuinely helps independently-splittable tasks (e.g., research) but Anthropic's internal eval showed it costs ~15x a normal chat interaction, and is less effective on tightly coupled tasks like coding where steps can't be parallelized. Use a capable lead with cheaper subagents to reduce the multiplier; failure handling multiplies with agent count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.6 Reliability Floor&lt;/strong&gt;&lt;br&gt;
Define a base (retry budget, latency ceiling) first, then tune cost above it, never below — cost pressure is more visible daily than reliability pressure, so the eval's pinned baseline score is what makes the floor enforceable before shipping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.5.7 Postmortem: Fan-Out on a Tightly-Coupled Task&lt;/strong&gt;&lt;br&gt;
Parallel fan-out applied to a sequential, dependent-step task tripled the bill with barely improved quality, since subagents mostly waited on each other. Reverting to single-agent restored cost with equal quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  4.6 Security
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.6.1 Prompt Injection: The Core Threat&lt;/strong&gt;&lt;br&gt;
The model processes its entire context as one undifferentiated token stream — no structural boundary separates trusted instructions from untrusted data, so hidden instructions in fetched content are read as commands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.2 Defense: Treat Content as Data&lt;/strong&gt;&lt;br&gt;
Treat all fetched/user-supplied content as data to examine, never instructions to follow — trusting your own users doesn't help, since the hostile instruction typically arrives via retrieved content. Delimiter-wrapping helps but is a soft boundary; the reliable boundary is what the agent is allowed to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.3 Threat Scope &amp;amp; Jailbreak vs. Injection&lt;/strong&gt;&lt;br&gt;
Any content someone else can write is a vector (shared docs, DB records, emails, chained tool outputs), direct or hidden. Jailbreaks target the model's own safety constraints; injections hijack the application's instructions — both need the same layered defense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.4 Secure-by-Design Identity and Access&lt;/strong&gt;&lt;br&gt;
Agent identity carries only the narrowest permission set the task requires; secrets live in env vars or a secret manager, never committed config. Least privilege bounds the blast radius of a successful injection; committed secrets are permanent exposures since only rotation (impossible for a hardcoded value) fixes a leak.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.5 Hook-Based Guardrails&lt;/strong&gt;&lt;br&gt;
A &lt;code&gt;PreToolUse&lt;/code&gt; hook blocks a tool call before execution and logs every privileged action — enforced control vs. an unenforced prompt-level rule. Precedence when rules conflict: deny &amp;gt; ask &amp;gt; allow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.6 Scoping for Regulated Review&lt;/strong&gt;&lt;br&gt;
Three early questions: where is data processed (residency), how is access logged (per-action audit trail), and can configuration be centrally administered (prevents developers quietly widening permissions).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.7 ZDR Note&lt;/strong&gt;&lt;br&gt;
Zero Data Retention eligibility varies by model and platform, not guaranteed even under an existing agreement — confirm against Anthropic's Trust Center (and Bedrock/Vertex/Foundry retention policies separately) at scoping time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.8 Layered Security Model&lt;/strong&gt;&lt;br&gt;
Model training + classifiers (reduce injection landing) → treating content as data (reduce acted-upon injections) → least privilege + locked config (bound blast radius) → hooks (enforce + record) → regulated-review scoping (make it auditable) — no single layer is sufficient alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.9 OS-Level Sandboxing&lt;/strong&gt;&lt;br&gt;
Isolates at the process level regardless of hook/identity config: filesystem isolation restricts the agent to its working directory, network isolation restricts outbound connections to a named endpoint set. Holds even when a hook is missing or bypassed — typically the first thing enterprise security reviewers ask about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.10 Postmortem: Hidden Instruction Redirected a Write&lt;/strong&gt;&lt;br&gt;
An internal-only agent skipped validating fetched web content, assuming trusting the user meant trusting the request; a hidden instruction in a fetched page redirected the agent's write to an unintended path. Fix: treat fetched content as data plus a &lt;code&gt;PreToolUse&lt;/code&gt; hook denying writes outside the permitted path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4.6.11 Defense Checklist&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Threat&lt;/th&gt;
&lt;th&gt;Entry Point&lt;/th&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;th&gt;Logged&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt injection&lt;/td&gt;
&lt;td&gt;Hidden instructions in fetched content&lt;/td&gt;
&lt;td&gt;Treat as data + hook refusing untrusted-triggered actions&lt;/td&gt;
&lt;td&gt;Source, attempted action, block&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jailbreak&lt;/td&gt;
&lt;td&gt;Crafted user prompt&lt;/td&gt;
&lt;td&gt;Input validation + model action constraints&lt;/td&gt;
&lt;td&gt;Flagged prompt, refusal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Over-broad access&lt;/td&gt;
&lt;td&gt;Identity scoped wider than needed&lt;/td&gt;
&lt;td&gt;Least privilege, secrets manager, locked auth config&lt;/td&gt;
&lt;td&gt;Every privileged action + identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sandbox escape&lt;/td&gt;
&lt;td&gt;Steered agent reaching uncovered paths/endpoints&lt;/td&gt;
&lt;td&gt;OS-level filesystem/network isolation&lt;/td&gt;
&lt;td&gt;Every denied access attempt + trigger&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  4.7 Cumulative Task
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;4.7.1 Cumulative Task Structure&lt;/strong&gt;&lt;br&gt;
A single runnable application containing three planted defects, one per layer (eval/testing, failure-handling, security) — e.g., an agent that writes based on untrusted fetched content without a hook, a retry loop with no real backoff that retries terminal errors too, and no &lt;code&gt;is_error&lt;/code&gt; handling on tool results.&lt;/p&gt;




&lt;h2&gt;
  
  
  Module 5: Accelerators &amp;amp; IP Contribution
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5.1 Module Introduction
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.1.1 Core Theme&lt;/strong&gt;&lt;br&gt;
A build that works is not yet a build that survives reuse, review, or deployment — templates aren't configurable, contributions aren't verifiable by a stranger, models aren't pinned, platforms haven't cleared compliance, and trust boundaries haven't been mapped. Much of this work is driven by the customer's cloud/compliance posture, not technical preference.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.2 Packaging for Reuse
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.2.1 Accelerator Definition&lt;/strong&gt;&lt;br&gt;
A solution packaged so future engagements start from a working foundation instead of a blank repo — separating engagement-specific code from a parameterized reusable core. Package while the build is fresh, before the reasoning behind hardcoded values is lost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.2.2 Three Asset Types&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Asset Type&lt;/th&gt;
&lt;th&gt;What It Bundles&lt;/th&gt;
&lt;th&gt;Correct Packaging Requires&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agent Template&lt;/td&gt;
&lt;td&gt;System prompt, tool schemas, loop structure&lt;/td&gt;
&lt;td&gt;Pull domain values into configuration with documented defaults — new team sets values, doesn't edit the loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP Server Package&lt;/td&gt;
&lt;td&gt;Exposed tools, their inputs, controllable scope&lt;/td&gt;
&lt;td&gt;Document each tool input; let the installing team set scope — installs without code edits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval Suite&lt;/td&gt;
&lt;td&gt;Graded test set + judge rubric&lt;/td&gt;
&lt;td&gt;Ship dataset and rubric together as the deployment gate (run against a pinned baseline before promoting a new model version)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.2.3 Common Failure: Loose Scripts&lt;/strong&gt;&lt;br&gt;
Shipping an agent as loose scripts instead of a template looks reusable because it "runs," but customer-specific values stay buried across files, so the next team copies and diverges instead of configuring one asset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.2.4 Documentation &amp;amp; Audit Bundling&lt;/strong&gt;&lt;br&gt;
Documentation must cover environment assumptions, expected inputs, handled failure modes, and the defining eval, or the next team treats the asset as a black box. Bundle the audit log (data touched, identity acted under) too — a regulated reviewer asks for this at the first security review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.2.5 Packaging Checklist&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Asset Type&lt;/th&gt;
&lt;th&gt;Parameterize&lt;/th&gt;
&lt;th&gt;Document&lt;/th&gt;
&lt;th&gt;Bundle for Audit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agent template&lt;/td&gt;
&lt;td&gt;Prompts, paths, scopes, credentials by reference, thresholds&lt;/td&gt;
&lt;td&gt;Environment assumptions, expected inputs, handled failures, defining eval&lt;/td&gt;
&lt;td&gt;Data touched, identity acted under, action log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP Server&lt;/td&gt;
&lt;td&gt;Scopes, credentials by reference, per-customer paths&lt;/td&gt;
&lt;td&gt;Expected inputs per tool, scope boundaries, handled failures&lt;/td&gt;
&lt;td&gt;Data touched, identity acted under, action log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval Suite&lt;/td&gt;
&lt;td&gt;Thresholds, dataset paths&lt;/td&gt;
&lt;td&gt;Rubric logic, what scores mean, pinned baseline&lt;/td&gt;
&lt;td&gt;Data touched, identity acted under, action log&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.2.6 Postmortem: Hardcoded Values Labeled Reusable&lt;/strong&gt;&lt;br&gt;
A team hardcoded customer-specific values into a template to hit a deadline, then labeled it "reusable"; a second team couldn't configure it (no parameters, no documentation, no bundled eval) and had to rewrite it fully. A template that runs has not been packaged for reuse — these are different finishing states.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.3 Contributing Back
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.3.1 Definition&lt;/strong&gt;&lt;br&gt;
Moving an asset from private reuse to shared infrastructure through a documented channel carrying version, install steps, and components as one unit — an asset already packaged for internal reuse is already close to what a maintainer needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.3.2 Matching Contribution to Channel&lt;/strong&gt;&lt;br&gt;
The Claude Cookbook takes self-contained, single/multi-pattern reference implementations demonstrated end to end, not a full application; open-source MCP servers/tools each have their own repo and conventions. Misplacement — sending a full multi-component app to the Cookbook — is one of the most common reasons a contribution never gets reviewed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.3.3 Four Things That Make Verification Possible&lt;/strong&gt;&lt;br&gt;
(1) Does one thing, (2) an example shows it running, (3) a test proves it works, (4) a short statement names the assumptions — the bar is set by what needs checking, not code cleverness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.3.4 Rights and Attribution Come First&lt;/strong&gt;&lt;br&gt;
Licensing decides whether a contribution can be accepted at all; code carried in from a customer engagement may have constraints on where it can go. Confirming the right to contribute and attributing prior work is a gate that must be passed before technical review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.3.5 Contribution-Readiness Reference&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Channel&lt;/th&gt;
&lt;th&gt;What a Maintainer Checks&lt;/th&gt;
&lt;th&gt;Licensing/Attribution&lt;/th&gt;
&lt;th&gt;Example/Test Bar&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cookbook (focused example) or tool/server's own repo&lt;/td&gt;
&lt;td&gt;Code does one thing, fully readable&lt;/td&gt;
&lt;td&gt;Confirm right to contribute engagement code, prior work attributed&lt;/td&gt;
&lt;td&gt;Runnable example + a test proving behavior, not just description&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.3.6 Postmortem: Unreviewed PR for Three Weeks&lt;/strong&gt;&lt;br&gt;
A PR sat unreviewed because it had no test, no example, and no stated environment assumptions — the maintainer couldn't verify it without reconstructing the developer's work. A contribution the reviewer can't verify sits at the back of the queue regardless of code correctness.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.4 Requirements &amp;amp; Lifecycle
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.4.1 From Business Problem to Functional Requirements&lt;/strong&gt;&lt;br&gt;
A functional requirement states what the system must do with enough detail to check (e.g., "classify each ticket into one of four queues... never auto-send without human approval," not "help agents answer faster") — a specific goal becomes an eval line and a review criterion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.4.2 Deriving Infrastructure Requirements&lt;/strong&gt;&lt;br&gt;
Non-functional constraints derived by asking: latency (how fast, measured where?), scale (how many requests, at what peak?), residency (where must data be processed?), identity (who acts, under what credentials, what's auditable?) — these four most often decide the deployment platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.4.3 Documenting Requirements&lt;/strong&gt;&lt;br&gt;
A short record of functional behaviors, infrastructure constraints, and the regulation each derives from lets a platform choice be defended as following from requirements rather than familiarity.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.5 Systems Lifecycle for Claude Applications
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.5.1 Seven Phases&lt;/strong&gt;&lt;br&gt;
Requirements (capture functional/infrastructure needs) → Design (platform, model, trust boundaries) → Build (agent, tools, prompts) → Test (evals, unit/integration/e2e) → Deploy (pin version, gate on eval) → Operate (instrument cost/latency/errors, enforce guardrails) → Iterate (feed production findings back into requirements).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.5.2 Gating Between Phases&lt;/strong&gt;&lt;br&gt;
A gate is the decision point to move between phases where a regulated engagement retains control (e.g., don't move design→build until the platform satisfies residency). Refusing to skip a gate is what keeps an application reviewable; a one-off experiment may collapse phases, a regulated deployment cannot.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.6 Deployment &amp;amp; Versioning
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.6.1 Platform Choice Driven by Customer's Cloud&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Identity/Data Model&lt;/th&gt;
&lt;th&gt;When to Choose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First-party Claude API&lt;/td&gt;
&lt;td&gt;Anthropic identity and terms&lt;/td&gt;
&lt;td&gt;No binding cloud/residency constraint; wants newest capabilities first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Platform on AWS&lt;/td&gt;
&lt;td&gt;Anthropic identity/terms via customer's AWS account; inference outside the AWS boundary&lt;/td&gt;
&lt;td&gt;On AWS but wants Anthropic model IDs/lifecycle parity with first-party API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude in Amazon Bedrock&lt;/td&gt;
&lt;td&gt;Messages API at &lt;code&gt;/anthropic/v1/messages&lt;/code&gt;; data stays inside customer's AWS boundary&lt;/td&gt;
&lt;td&gt;On AWS, wants feature parity + compliance posture there&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude on Amazon Bedrock (legacy)&lt;/td&gt;
&lt;td&gt;AWS identity/billing; &lt;code&gt;InvokeModel&lt;/code&gt;/&lt;code&gt;Converse&lt;/code&gt; APIs, ARN-versioned IDs&lt;/td&gt;
&lt;td&gt;Existing (unmigrated) Bedrock integration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Vertex AI&lt;/td&gt;
&lt;td&gt;Google Cloud identity/IAM/billing; regional or global endpoints&lt;/td&gt;
&lt;td&gt;On Google Cloud with compliance posture there&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Third-party (e.g., Microsoft Foundry)&lt;/td&gt;
&lt;td&gt;Wrapping product's identity/billing&lt;/td&gt;
&lt;td&gt;Already runs the platform embedding Claude; residency depends on hosting form (Azure-hosted vs. Anthropic-hosted) per model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.6.2 Identity and Residency Answered by Platform&lt;/strong&gt;&lt;br&gt;
Bedrock uses AWS identity and keeps data in the customer's AWS boundary; Vertex uses Google Cloud identity/boundary; both offer regional routing. Matching platform to the customer's existing compliance agreement avoids a residency review from scratch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.6.3 Pinning Versions&lt;/strong&gt;&lt;br&gt;
Every model ID points to a specific snapshot; aliases (e.g., "Opus," "Sonnet") evolve over time, so pin the full model ID, not the alias (e.g., &lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt; vs. the moving &lt;code&gt;claude-haiku-4-5&lt;/code&gt;), and version prompts/assets alongside code with a rollback-ready prior version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.6.4 Promote via the Eval&lt;/strong&gt;&lt;br&gt;
Send a new version to a portion of traffic, compare against the pinned baseline, promote or roll back on the result — the eval is the deployment gate, not just a one-time test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.6.5 Deployment-Platform Versioning&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;Versioning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First-party API&lt;/td&gt;
&lt;td&gt;Pin full model ID, keep prior snapshot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Platform on AWS&lt;/td&gt;
&lt;td&gt;Same ID format as Claude API; lifecycle follows Anthropic's schedule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude in Amazon Bedrock&lt;/td&gt;
&lt;td&gt;Pin full model ID with &lt;code&gt;anthropic.&lt;/code&gt; prefix; partner retirement dates differ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude on Amazon Bedrock (legacy)&lt;/td&gt;
&lt;td&gt;Pin via ARN-versioned identifiers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Vertex AI&lt;/td&gt;
&lt;td&gt;Pin full model ID before rollout; partner retirement dates differ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Third-party platform&lt;/td&gt;
&lt;td&gt;Pin per the platform's own versioning controls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.6.6 Postmortem: Moving Alias Broke Production&lt;/strong&gt;&lt;br&gt;
A deployment shipped against a moving alias ("opus"); it silently advanced to a new version, breaking downstream parsing with no pinned prior version to roll back to, forcing a hotfix instead. Lesson: pin the full model ID, retain the prior pinned version, gate promotions through the eval.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.7 Comparing Platforms
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.7.1 Latency&lt;/strong&gt;&lt;br&gt;
Depends on platform location relative to the customer and feature-access timing (first-party API usually gets new capabilities first) — must be measured from the customer's actual region/payload. Within Bedrock, global vs. regional endpoints is the primary residency control and can affect cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.7.2 Compliance Often Ends the Debate&lt;/strong&gt;&lt;br&gt;
A customer already certified on one cloud is unlikely to re-certify on another; residency/certifications/audit access differ by platform and are pass-or-fail for regulated customers. First-party API may lack EU residency (Bedrock/Vertex typically required); on Foundry, hosting is per-model. Raise compliance constraints at scoping, not at contract review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.7.3 Cost Beyond Per-Token Rate&lt;/strong&gt;&lt;br&gt;
Token rates are broadly aligned across platforms; total cost is driven by egress, platform fees, and integration effort — instrument cost per call per platform rather than comparing token price alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.7.4 Cross-Platform Comparison Reference&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;How It Differs&lt;/th&gt;
&lt;th&gt;How to Measure&lt;/th&gt;
&lt;th&gt;Winner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;In-region platform shortens round trip; first-party API gets features first&lt;/td&gt;
&lt;td&gt;From customer's actual region + payload&lt;/td&gt;
&lt;td&gt;In-region cloud wins on latency; first-party wins on feature access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compliance&lt;/td&gt;
&lt;td&gt;Residency, certifications, audit controls vary by platform&lt;/td&gt;
&lt;td&gt;Against customer's existing certification/residency requirements at scoping&lt;/td&gt;
&lt;td&gt;The already-certified platform wins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Token price, egress, platform fees, integration effort all vary&lt;/td&gt;
&lt;td&gt;Total cost per call including egress/integration&lt;/td&gt;
&lt;td&gt;Lowest total cost for the actual workload&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.7.5 Postmortem: Familiar Platform Failed Residency Review&lt;/strong&gt;&lt;br&gt;
A team picked the platform they knew best for a regulated customer since migration was fast; it passed functional tests but failed the customer's residency requirement at security review, requiring a rebuild. Familiarity optimizes for build speed, not for whether the deployment is allowed to ship.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.8 Trust Boundaries
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.8.1 Multi-Component Coordination&lt;/strong&gt;&lt;br&gt;
An app might chain an API request → Claude Code task → MCP server reaching a customer system; each connection creates a place where identity, secrets, and untrusted input can cross — map what each component does before connecting anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.8.2 Least Privilege Applies to the Whole Application&lt;/strong&gt;&lt;br&gt;
Each component operates under its own identity, scoped to only what its task needs — the application is only as contained as its most privileged seam, so one overly-broad component becomes the weak point even if others are properly scoped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.8.3 Regulated Review Requirements&lt;/strong&gt;&lt;br&gt;
Requires justifying audit logging, data-residency decisions, and permission controls across the full application, not per component — confirm ZDR/HIPAA BAA eligibility for each individual component against the Trust Center before scoping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5.8.4 Multi-Component Integration Map&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Contributes&lt;/th&gt;
&lt;th&gt;Trust Boundary at Seam&lt;/th&gt;
&lt;th&gt;Control&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First-party API&lt;/td&gt;
&lt;td&gt;Orchestrates workflow, holds entry point&lt;/td&gt;
&lt;td&gt;Request entering the app from outside&lt;/td&gt;
&lt;td&gt;Input validation + identity the call runs under&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code task&lt;/td&gt;
&lt;td&gt;Runs agentic work, may fetch external content&lt;/td&gt;
&lt;td&gt;Content it fetched — untrusted downstream&lt;/td&gt;
&lt;td&gt;Treat fetched content as data at the next seam&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP server&lt;/td&gt;
&lt;td&gt;Reaches a customer system to read/act&lt;/td&gt;
&lt;td&gt;System access held on the app's behalf&lt;/td&gt;
&lt;td&gt;Scope to least privilege + log the access&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.8.5 Postmortem: Untrusted Content Passed as Trusted Input&lt;/strong&gt;&lt;br&gt;
Three components each passed their own tests; content fetched by the Claude Code task was passed directly into the next call as trusted input with no boundary control at that seam — a hidden instruction there would have executed. A component being trusted in isolation says nothing about the seam leaving it.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.9 Cumulative Task
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;5.9.1 Cumulative Task Structure&lt;/strong&gt;&lt;br&gt;
A single deployed accelerator with three planted defects: a hardcoded customer-specific value where a parameter belongs (packaging), a moving model alias instead of a pinned full ID (versioning), and fetched untrusted content passed directly into the next call with no boundary control (trust boundary). Task: identify all three, explain the runtime consequence, and write the corrected lines.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>learning</category>
      <category>llm</category>
    </item>
    <item>
      <title>Everything I Learned From My Claude Code Training</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Mon, 31 Aug 2026 17:28:38 +0000</pubDate>
      <link>https://dev.to/yashnigam/everything-i-learned-from-my-claude-code-training-1k94</link>
      <guid>https://dev.to/yashnigam/everything-i-learned-from-my-claude-code-training-1k94</guid>
      <description>&lt;h1&gt;
  
  
  Everything I Learned From My Claude Code Training
&lt;/h1&gt;

&lt;p&gt;I recently went through a hands-on training on &lt;strong&gt;Claude Code&lt;/strong&gt; and the broader Claude ecosystem. Below are my complete notes, organized by topic, with screenshots from the session included at each relevant step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
The Claude Ecosystem

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;AI-Code-Design-Cowork-Security-Chrome-Office&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
The Agentic Loop (Query Engine)

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Architecture&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Models, Effort, and Slash Commands

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Models-Commands&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Session Management

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Export-Fork-Resume-Compact&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Project Configuration

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;ClaudeFolder-Permissions-Init-ClaudeMd&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Auto Mode vs. Plan Mode

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Auto-Plan&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Commands vs. Skills vs. Agents

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Standup-Metaprompting-PR-Invocation&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Code Review via Subagents

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Security-Performance-Coverage-Quality-Usage-Deploy-Setup-Execution-Tests-Stack-Creation-Example-Focus-OWASP-Memory-Output-Background&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Multi-Agent Harness Concepts

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;SecondPass-Comparison&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Hooks

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Flow-PreToolUse-Types-Blocking-Slack-Sharing&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
MCP (Model Context Protocol)

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Concept-Rationale-Types&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
End-to-End Automation

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;End-to-end automation of security review → issue creation → issue fix → PR creation → PR fix, all through agents.&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Agent Teams

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Multiple agents (security, performance, test) talk to each other and take actions autonomously — filing issues, creating PRs, merging PRs, and closing issues, all automated.&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Key Takeaways

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Loop-Config-Commands-Hooks-Sessions-Teams&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  1. The Claude Ecosystem
&lt;/h2&gt;

&lt;p&gt;Claude Code is one tool within a larger ecosystem, not a standalone product:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Claude AI&lt;/strong&gt; — the conversational assistant&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; — the agentic coding CLI (the focus of this article), which has:

&lt;ul&gt;
&lt;li&gt;An &lt;strong&gt;agentic loop&lt;/strong&gt; that loops until the task is completed&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;tool system&lt;/strong&gt;: git, bash, read, write files&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;permission pipeline&lt;/strong&gt;, configurable per tool (Allow / Deny / Ask)&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;context manager&lt;/strong&gt;: compaction and pruning of context within a session&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session persistence&lt;/strong&gt;: work can be paused and resumed later, with context and memory maintained&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Design&lt;/strong&gt; — design tooling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Co-work&lt;/strong&gt; — collaboration tooling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Security&lt;/strong&gt; — security-focused tooling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Chrome plugin&lt;/strong&gt; — browser automation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Office plugin&lt;/strong&gt; — office document integration&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In the backend, Claude Code talks to the LLM hosted by Anthropic via the &lt;code&gt;/v1/messages&lt;/code&gt; API.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram of the agentic loop: gather context, take action, verify results, repeat until done.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3wtmefqtym6qcz2k6pd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3wtmefqtym6qcz2k6pd.png" alt="Claude ecosystem overview" width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Tools like Cursor, Windsurf, and Perplexity act as &lt;strong&gt;integrators&lt;/strong&gt; — giving access to multiple LLMs. Claude, by contrast, offers an entire vertically integrated ecosystem.&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h2&gt;
  
  
  2. The Agentic Loop (Query Engine)
&lt;/h2&gt;

&lt;p&gt;This is the core engine behind Claude Code. Given a request — e.g. &lt;em&gt;"Create a unit test case for my function"&lt;/em&gt; — the loop works like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Gather context&lt;/strong&gt; — goes to the file, reads the imported files also, reads the function in the file (which needs to be mocked).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send context to the LLM&lt;/strong&gt; — the LLM decides, based on the context, what unit test cases must be created after it has all the info, and tells Claude Code what actions to take (e.g. "create new test case").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use tools&lt;/strong&gt; — Claude Code creates or modifies the file using its tool system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify&lt;/strong&gt; — the agent checks whether the job is done or not (e.g. runs the test to confirm it passes).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Additional properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;user can modify the context&lt;/strong&gt; in between the loop or after it.&lt;/li&gt;
&lt;li&gt;The loop &lt;strong&gt;runs continuously&lt;/strong&gt; until the task is not done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything depends on the LLM&lt;/strong&gt; being used — so it's worth knowing which LLM is best suited for the task.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Agentic Coding Tools Architecture
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Three-layer architecture diagram showing coding agents (Claude Code, Copilot, Codex, Cursor, etc.) between the chat interface and various LLM hosting backends.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F42xif6ull04n8qh8ldsg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F42xif6ull04n8qh8ldsg.png" alt="Agentic coding tools architecture diagram 1" width="800" height="435"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude Code's runtime architecture, from terminal command through the CLI parser and QueryEngine agentic loop to tool calls and API routing.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftktqmbdpi8agxcsqsna0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftktqmbdpi8agxcsqsna0.png" alt="Agentic coding tools architecture diagram 2" width="800" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code always connects to the Anthropic base URL.&lt;/li&gt;
&lt;li&gt;It can also connect to &lt;strong&gt;Ollama&lt;/strong&gt; and open-source models such as &lt;strong&gt;Qwen&lt;/strong&gt; or &lt;strong&gt;Llama 3.2&lt;/strong&gt; — enabling a fully local setup. This is currently the only fully local option available.&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  3. Models, Effort, and Slash Commands
&lt;/h2&gt;

&lt;p&gt;Models, effort levels, and sessions are the main levers for controlling how Claude Code works.&lt;/p&gt;
&lt;h3&gt;
  
  
  Models
&lt;/h3&gt;

&lt;p&gt;Four main models, chosen based on the type of task and usage:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fable&lt;/strong&gt; — for complex tasks: architecture, design, security, coding&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Opus&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sonnet&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Haiku&lt;/strong&gt; — for simpler tasks&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Qwen-local&lt;/code&gt; is the local-model equivalent of Haiku.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;opusplan&lt;/code&gt; = Opus for planning + Sonnet for doing the task.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Guidance:&lt;/strong&gt; try a smaller model at higher effort first — otherwise switch to a new/bigger model as required. Use Fable only if Opus cannot handle the task.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Terminal status line showing the active model, session name, context usage, token count, and running cost.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffegodowq5s12f56s5oub.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffegodowq5s12f56s5oub.png" alt="Models and effort screen clipping" width="657" height="116"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Screen clippings taken: 29-08-2026 02:21 and 29-08-2026 02:14&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Slash commands for managing everything
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;/model&lt;/code&gt; — switch between different models&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/effort&lt;/code&gt; — decides the effort to put in: &lt;strong&gt;high → xhigh → max&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/statusline&lt;/code&gt; — shown in the Claude prompt, displays total tokens used; helps decide when to compact the session&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  4. Session Management
&lt;/h2&gt;

&lt;p&gt;Before starting with Claude Code, it's worth understanding &lt;code&gt;/sessions&lt;/code&gt; — session management.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;We can &lt;strong&gt;resume from wherever we stopped&lt;/strong&gt; using a session ID.&lt;/li&gt;
&lt;li&gt;Start a new session: &lt;code&gt;claude -n "session_name"&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Resume a previous session: &lt;code&gt;claude --resume&lt;/code&gt;, &lt;code&gt;claude --continue&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/status&lt;/code&gt; — check session status&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Practical guidance
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;We can create &lt;strong&gt;3 different sessions&lt;/strong&gt; for 3 different issues, potentially working on 3 different models.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/export explore-current-project.md&lt;/code&gt; — exports the conversation to a markdown file, e.g.:
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   ⎿  Conversation exported to: H:\Claude-Code-Projects-17-19-Aug\claude_showcase_17_aug_26\explore-current-project.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;/fork&lt;/code&gt; — forks the current session so you can go to a new forked session without modifying the original session.&lt;/li&gt;
&lt;li&gt;If &lt;strong&gt;80% of context&lt;/strong&gt; is used, we can compact it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/resume&lt;/code&gt; — switch between sessions.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/compact&lt;/code&gt; — summarizes the session data; some data could be removed in the process.&lt;/li&gt;
&lt;li&gt;If old context is not found in the current session, tokens would need to be spent again to recreate it — and &lt;strong&gt;wrong information could get compacted&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;There is a &lt;strong&gt;usage limit at the account level&lt;/strong&gt;; every session uses a part of that limit.&lt;/li&gt;
&lt;li&gt;We should create multiple sessions — e.g. 3 different sessions for 3 different bugs. Information for the first bug may not be needed for the second bug, and sessions help Claude read old information without wasting tokens to recreate it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Claude Code's /status panel showing version, session details, login/account info, and connected MCP servers.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpsn4bh63mn1aq0u7bkn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpsn4bh63mn1aq0u7bkn.png" alt="Session export screen clipping" width="689" height="318"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  5. Project Configuration
&lt;/h2&gt;
&lt;h3&gt;
  
  
  README.md
&lt;/h3&gt;

&lt;p&gt;The most important file for reading all the information about a project.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Before telling the tool what to do, plan first&lt;/strong&gt; — create the architecture first.&lt;/li&gt;
&lt;li&gt;Decide on a design template.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  CLAUDE.md — Rules
&lt;/h3&gt;

&lt;p&gt;Used to control the behaviour of Claude:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which JavaScript library to use.&lt;/li&gt;
&lt;li&gt;Preventing Claude Code from reading sensitive files so it does not use them.&lt;/li&gt;
&lt;li&gt;How to give Claude permission to access your code — what it can read and what it cannot read.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  &lt;code&gt;.claude&lt;/code&gt; folder and settings
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;First, create a &lt;code&gt;.claude&lt;/code&gt; folder.&lt;/li&gt;
&lt;li&gt;Create &lt;code&gt;settings.json&lt;/code&gt; — accessible to the whole team, pushed to the repo.&lt;/li&gt;
&lt;li&gt;Create &lt;code&gt;settings.local.json&lt;/code&gt; — will &lt;strong&gt;not&lt;/strong&gt; be pushed to the repo, and will override &lt;code&gt;settings.json&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here you can decide which model to use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"opusplan"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Permissions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permissions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm run *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git status)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git diff)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git add *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git commit *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Write(src/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Write(tests/**)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ask"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git push:*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm install *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Write(package.json)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(./.env*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(./secrets/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(./**/credentials*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rm -rf:*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(curl:*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(wget:*)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Claude declining to read a blocked .env.example file per project settings and offering alternative ways to proceed.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gmdbkyd5x4jg3llkk6f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gmdbkyd5x4jg3llkk6f.png" alt="Settings configuration screen clipping" width="800" height="115"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  &lt;code&gt;/init&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Initializes a &lt;code&gt;CLAUDE.md&lt;/code&gt; file with codebase documentation after going through the project at the root level.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add conditions for Claude to always follow — should not go above &lt;strong&gt;150 to 200 lines max&lt;/strong&gt;; this is very specific to the root of the project.&lt;/li&gt;
&lt;li&gt;We can have a &lt;code&gt;CLAUDE.md&lt;/code&gt; file for &lt;strong&gt;every sub-folder&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;If there is no &lt;code&gt;CLAUDE.md&lt;/code&gt; file, don't rely on Claude having project-specific context — use &lt;code&gt;CLAUDE.md&lt;/code&gt; for every project.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Example: the actual CLAUDE.md used in this training repo
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Coding Standards&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Always use async/await, never raw Promises
&lt;span class="p"&gt;-&lt;/span&gt; All functions must have JSDoc comments
&lt;span class="p"&gt;-&lt;/span&gt; No console.log in production code — use the logger utility
&lt;span class="p"&gt;-&lt;/span&gt; Every new function must have at least one unit test
&lt;span class="p"&gt;-&lt;/span&gt; Pure validation functions use a single return statement with a composed expression
&lt;span class="p"&gt;-&lt;/span&gt; Never use multiple early-return guard clauses in validator functions

&lt;span class="gu"&gt;## Repository conventions&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; CommonJS (&lt;span class="sb"&gt;`require`&lt;/span&gt;/&lt;span class="sb"&gt;`module.exports`&lt;/span&gt;) throughout, not ESM.
&lt;span class="p"&gt;-&lt;/span&gt; ESLint enforces: &lt;span class="sb"&gt;`eqeqeq`&lt;/span&gt; (always &lt;span class="sb"&gt;`===`&lt;/span&gt;), &lt;span class="sb"&gt;`curly`&lt;/span&gt;, &lt;span class="sb"&gt;`no-var`&lt;/span&gt;/&lt;span class="sb"&gt;`prefer-const`&lt;/span&gt;,
  &lt;span class="sb"&gt;`no-return-await`&lt;/span&gt;, &lt;span class="sb"&gt;`require-await`&lt;/span&gt; (no &lt;span class="sb"&gt;`async`&lt;/span&gt; functions without an &lt;span class="sb"&gt;`await`&lt;/span&gt;),
  unused-arg exception for &lt;span class="sb"&gt;`_`&lt;/span&gt;-prefixed names.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`.claude/settings.json`&lt;/span&gt; denies reads of &lt;span class="sb"&gt;`.env*`&lt;/span&gt;, &lt;span class="sb"&gt;`secrets/**`&lt;/span&gt;, and
  &lt;span class="sb"&gt;`credentials*`&lt;/span&gt; — don't try to work around this to inspect real secrets.

&lt;span class="gu"&gt;## What Claude Must Never Do&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Never modify .env or .env.&lt;span class="se"&gt;\*&lt;/span&gt; files
&lt;span class="p"&gt;-&lt;/span&gt; Never push directly to main branch
&lt;span class="p"&gt;-&lt;/span&gt; Never remove existing tests
&lt;span class="p"&gt;-&lt;/span&gt; Never install packages without confirming with the developer

&lt;span class="gu"&gt;## PR and Git Standards&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Commit messages follow Conventional Commits: feat:, fix:, docs:, test:
&lt;span class="p"&gt;-&lt;/span&gt; PR descriptions must include: what changed, why it changed, how to test

&lt;span class="gu"&gt;## Tech Stack&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Runtime: Node.js 20
&lt;span class="p"&gt;-&lt;/span&gt; Framework: Express.js
&lt;span class="p"&gt;-&lt;/span&gt; Database: PostgreSQL with Prisma ORM
&lt;span class="p"&gt;-&lt;/span&gt; Testing: Jest
&lt;span class="p"&gt;-&lt;/span&gt; Language: ECMAScript 2022 for all new code only

&lt;span class="gu"&gt;## Architecture&lt;/span&gt;
Request flow: &lt;span class="sb"&gt;`src/index.js`&lt;/span&gt; wires Express with &lt;span class="sb"&gt;`requestLogger`&lt;/span&gt; middleware
globally, mounts all auth endpoints under &lt;span class="sb"&gt;`/api/auth`&lt;/span&gt; from &lt;span class="sb"&gt;`src/api/routes.js`&lt;/span&gt;,
and registers &lt;span class="sb"&gt;`errorHandler`&lt;/span&gt; last as the global error middleware.

Layering:
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/api/routes.js`&lt;/span&gt; — route handlers only; validates input shape, calls into
  &lt;span class="sb"&gt;`authService`&lt;/span&gt;, maps thrown errors to HTTP status codes (e.g. &lt;span class="sb"&gt;`'Invalid
  credentials'`&lt;/span&gt; → 401). It does NOT talk to &lt;span class="sb"&gt;`tokenHelper.js`&lt;/span&gt; directly.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/api/middleware.js`&lt;/span&gt; — cross-cutting concerns: &lt;span class="sb"&gt;`requestLogger`&lt;/span&gt;,
  &lt;span class="sb"&gt;`authenticate`&lt;/span&gt; (verifies Bearer token via &lt;span class="sb"&gt;`tokenHelper.verifyToken`&lt;/span&gt;, attaches
  &lt;span class="sb"&gt;`req.user`&lt;/span&gt;), &lt;span class="sb"&gt;`validateBody`&lt;/span&gt; (checks required fields present), &lt;span class="sb"&gt;`errorHandler`&lt;/span&gt;
  (catch-all, returns generic 500).
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/auth/authService.js`&lt;/span&gt; — business logic: &lt;span class="sb"&gt;`loginUser`&lt;/span&gt;, &lt;span class="sb"&gt;`refreshToken`&lt;/span&gt;,
  &lt;span class="sb"&gt;`revokeToken`&lt;/span&gt;, token generation. Owns the in-memory &lt;span class="sb"&gt;`refreshTokenStore`&lt;/span&gt; (a
  &lt;span class="sb"&gt;`Map`&lt;/span&gt;, standing in for a database — resets on every restart, not shared
  across processes).
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/auth/tokenHelper.js`&lt;/span&gt; — low-level JWT primitives: &lt;span class="sb"&gt;`verifyToken`&lt;/span&gt;
  (signature + expiry check, throws typed errors), &lt;span class="sb"&gt;`decodeToken`&lt;/span&gt; (no
  verification, for reading claims only), &lt;span class="sb"&gt;`extractBearerToken`&lt;/span&gt;, &lt;span class="sb"&gt;`getTokenTTL`&lt;/span&gt;.
  &lt;span class="sb"&gt;`authService.js`&lt;/span&gt; and &lt;span class="sb"&gt;`middleware.js`&lt;/span&gt; both depend on this; it depends on
  nothing else in the app.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/utils/validators.js`&lt;/span&gt; — pure, stateless input validators (email,
  password, UUID, sanitization). No I/O.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`src/utils/logger.js`&lt;/span&gt; — minimal structured JSON logger (stdout/stderr),
  level-gated by &lt;span class="sb"&gt;`LOG_LEVEL`&lt;/span&gt; env var. Not a real observability stack.

There's no real user database: &lt;span class="sb"&gt;`routes.js`&lt;/span&gt; builds a &lt;span class="sb"&gt;`mockUserRecord`&lt;/span&gt; inline
for login, and &lt;span class="sb"&gt;`authService.getUserById`&lt;/span&gt; returns a hardcoded stub — expect
these to be replaced with real persistence rather than extended in place.

JWT config (&lt;span class="sb"&gt;`JWT_SECRET`&lt;/span&gt;, &lt;span class="sb"&gt;`JWT_EXPIRES_IN`&lt;/span&gt;, &lt;span class="sb"&gt;`REFRESH_EXPIRES_IN`&lt;/span&gt;) is read from
env vars in both &lt;span class="sb"&gt;`authService.js`&lt;/span&gt; and &lt;span class="sb"&gt;`tokenHelper.js`&lt;/span&gt; independently, each
with its own fallback default — keep them in sync if changing defaults.

&lt;span class="gu"&gt;### Request flow (login example)&lt;/span&gt;
&lt;span class="sb"&gt;`POST /api/auth/login`&lt;/span&gt; → &lt;span class="sb"&gt;`validateBody(['email','password'])`&lt;/span&gt; → route handler
validates format via &lt;span class="sb"&gt;`validators.js`&lt;/span&gt; → &lt;span class="sb"&gt;`authService.loginUser()`&lt;/span&gt; checks bcrypt
hash, calls &lt;span class="sb"&gt;`generateAccessToken`&lt;/span&gt; (signed JWT) + &lt;span class="sb"&gt;`generateRefreshToken`&lt;/span&gt; (uuid
stored in &lt;span class="sb"&gt;`refreshTokenStore`&lt;/span&gt;) → returns &lt;span class="sb"&gt;`{ accessToken, refreshToken, user }`&lt;/span&gt;.
Protected routes (&lt;span class="sb"&gt;`/logout`&lt;/span&gt;, &lt;span class="sb"&gt;`/me`&lt;/span&gt;) go through &lt;span class="sb"&gt;`authenticate`&lt;/span&gt; middleware,
which calls &lt;span class="sb"&gt;`tokenHelper.verifyToken`&lt;/span&gt; and attaches the decoded payload to
&lt;span class="sb"&gt;`req.user`&lt;/span&gt;.

&lt;span class="gu"&gt;### Error convention&lt;/span&gt;
Route handlers translate known domain errors to HTTP status codes inline
(e.g. &lt;span class="sb"&gt;`'Invalid credentials'`&lt;/span&gt; → 401) and pass everything else to &lt;span class="sb"&gt;`next(err)`&lt;/span&gt;,
where &lt;span class="sb"&gt;`errorHandler`&lt;/span&gt; logs it and returns a generic 500. Preserve this pattern
when adding routes — don't leak internal error messages to clients from the
global handler.

&lt;span class="gu"&gt;### Token model&lt;/span&gt;
Two-token scheme: short-lived signed JWT access token (claims: &lt;span class="sb"&gt;`sub`&lt;/span&gt;, &lt;span class="sb"&gt;`email`&lt;/span&gt;,
&lt;span class="sb"&gt;`role`&lt;/span&gt;) + opaque UUID refresh token stored server-side in &lt;span class="sb"&gt;`refreshTokenStore`&lt;/span&gt;
with &lt;span class="sb"&gt;`{ userId, createdAt }`&lt;/span&gt;. Refresh tokens are revoked by deleting the map
entry (&lt;span class="sb"&gt;`revokeToken`&lt;/span&gt;, logout).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Example: Generating &lt;code&gt;validateCreditCard&lt;/code&gt; and Its Test Case from a Prompt
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: creating a function like this automatically creates a test case also (via the write-tests skill — covered in section 7).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Claude adding a JSDoc-documented validateCreditCard function (Luhn checksum, 13-19 digits) to validators.js.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7mmjqtr6ir0v1fp8qeyy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7mmjqtr6ir0v1fp8qeyy.png" alt="CLAUDE.md example screen clipping 1" width="800" height="506"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude adding Jest test cases for validateCreditCard covering valid Visa and Mastercard test numbers.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9yq8xcx9gf0uhgdrxj9j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9yq8xcx9gf0uhgdrxj9j.png" alt="CLAUDE.md example screen clipping 2" width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  6. Auto Mode vs. Plan Mode
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Auto mode&lt;/strong&gt; — Claude decides what to do automatically and does not ask for confirmation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan mode&lt;/strong&gt; — we can ask Claude to create a plan only, without executing.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;btw&lt;/code&gt; to ask a question which runs separately from the main section.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always switch to manual mode before making changes&lt;/strong&gt; you want to review first.&lt;/li&gt;
&lt;li&gt;In plan mode, Claude will ask you to go ahead with the changes first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Claude's plan mode proposal for a new validateBirthDate validator, with scope and constraints confirmed before execution.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5s82x6gw97xqs8wvpgy7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5s82x6gw97xqs8wvpgy7.png" alt="Plan mode screen clipping" width="800" height="363"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In auto mode, this behaviour changes — Claude proceeds independently.&lt;/p&gt;


&lt;h2&gt;
  
  
  7. Commands vs. Skills vs. Agents
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/commands&lt;/code&gt;&lt;/strong&gt; — executed &lt;strong&gt;manually&lt;/strong&gt;, defined via a markdown file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills&lt;/strong&gt; — run &lt;strong&gt;automatically&lt;/strong&gt; for us, based on what's asked. A command can be converted into a skill. Skills go into a skills folder.

&lt;ul&gt;
&lt;li&gt;There are a lot of third-party skills available.&lt;/li&gt;
&lt;li&gt;It's better to create our own skills, which use fewer tokens compared to those available externally, which tend to use more.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents&lt;/strong&gt; — specialized assistants that do a task in the background and come up with an answer. Run an agent separately in the background, or run it via a skill — a skill can also hand over a task to an agent.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Rule of thumb:&lt;/strong&gt; never use one-liners to do your job — don't go back to the tool again and again.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Example: custom command — &lt;code&gt;/standup&lt;/code&gt;
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;standup&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Generate a daily standup update from git history and open files&lt;/span&gt;
&lt;span class="na"&gt;disable-model-invocation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
You are a senior developer preparing a daily standup update for your team.
This is a Node.js/Express authentication API project. You have been working
in this codebase today. Check git log --since="00:00" --oneline and
git diff HEAD to understand what actually changed.

Draft a standup update based only on what you find in the git history
and open files — do not invent or assume work that is not visible.

Format:
Yesterday: [completed items, specific function or file names]
Today:     [in-progress items based on uncommitted changes or open TODOs]
Blockers:  [failing tests, TODO/FIXME comments, incomplete functions]

Keep each section to 3 bullet points maximum.
Use specific names — file names, function names, issue numbers where visible.
Be factual and brief. No filler. Write it as if reading it aloud in 30 seconds.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;em&gt;Output of the /standup command summarizing yesterday's commits, today's work, and current blockers.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmezi1ym8nuq1oz806ako.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmezi1ym8nuq1oz806ako.png" alt="Standup command screen clipping" width="800" height="219"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If Claude Code is used to create code, it should also be used to create the commit.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Metaprompting
&lt;/h3&gt;

&lt;p&gt;Take help of Claude to create a prompt using the &lt;strong&gt;RCTFCF&lt;/strong&gt; framework (Role, Context, Task, Constraint, Format):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Act as a prompt expert and help me create a structured prompt using role, context, task, constraint, and output format which can be used to do a security review for a Node.js API written in the Express framework. I want to check top 10 OWASP issues."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Also relevant: &lt;strong&gt;chain of thought prompting&lt;/strong&gt;, and the observation that &lt;strong&gt;LLMs understand prompting better in their own language&lt;/strong&gt; (i.e. structured, explicit formats they're trained to parse).&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude implementing validateDateOfBirth in validators.js following the ISO-date, no-future, 120-year-limit convention.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1bzyp8hd0rxxmivri0jc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1bzyp8hd0rxxmivri0jc.png" alt="Metaprompting screen clipping 1" width="800" height="230"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The /raise-pr command refusing to push directly to main and creating a feature branch instead.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fabndbnnfhjnhq018qhgn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fabndbnnfhjnhq018qhgn.png" alt="Metaprompting screen clipping 2" width="799" height="397"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Example: custom command — &lt;code&gt;/raise-pr&lt;/code&gt;
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;raise-pr&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Raise a GitHub pull request for the current branch targeting main&lt;/span&gt;
&lt;span class="na"&gt;disable-model-invocation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
You are a senior developer raising a pull request for a fix or feature branch.

Steps:
&lt;span class="p"&gt;1.&lt;/span&gt; Run git branch --show-current to get the current branch name.
&lt;span class="p"&gt;2.&lt;/span&gt; Run git log main..HEAD --oneline to list all commits on this branch.
&lt;span class="p"&gt;3.&lt;/span&gt; Run git diff main...HEAD to inspect all changes introduced by this branch.
&lt;span class="p"&gt;4.&lt;/span&gt; Check if the branch name or commits reference an issue number (e.g. fix/issue-204
   or "closes #204"). Extract it if present.

Using what you observe, generate a PR title and body:

Title format:
  type(scope): short summary under 72 characters  (same as the commit message)

Body format:
  ## What changed
&lt;span class="p"&gt;  -&lt;/span&gt; Bullet points describing each logical change
  ## Why it changed
&lt;span class="p"&gt;  -&lt;/span&gt; The problem or issue this PR resolves (reference issue number if found, e.g. closes #204)
  ## How to test
&lt;span class="p"&gt;  -&lt;/span&gt; Step-by-step instructions to verify the fix or feature works

Then run:
  gh pr create --base main --title "&lt;span class="nt"&gt;&amp;lt;title&amp;gt;&lt;/span&gt;" --body "&lt;span class="nt"&gt;&amp;lt;body&amp;gt;&lt;/span&gt;"

If the branch has no upstream yet, push it first:
  git push -u origin &lt;span class="nt"&gt;&amp;lt;branch-name&amp;gt;&lt;/span&gt;

Output only the PR title and body. No explanation, no commentary, no preamble.
Be precise and factual. Every word should earn its place.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Running &lt;code&gt;/raise-pr&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Permission prompt confirming the git push of the new feature branch to origin.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ujjq0qlugbz9x6ut8ut.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ujjq0qlugbz9x6ut8ut.png" alt="Raise PR screen clipping 1" width="652" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Confirmation that the feature branch was pushed and PR #1 was opened on GitHub.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4s16dp9snt8nnqfsnsy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4s16dp9snt8nnqfsnsy.png" alt="Raise PR screen clipping 2" width="800" height="204"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We can see the PR is created:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;GitHub pull requests tab showing the newly opened validateDateOfBirth PR.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffioz60p0f52ljnlhlhbi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffioz60p0f52ljnlhlhbi.png" alt="PR created screen clipping" width="800" height="227"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Skills are auto-invoked
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/commands&lt;/code&gt; are executed manually, but &lt;strong&gt;skills&lt;/strong&gt; get applied automatically based on the prompt. Seeing a list of all skills:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The /skills menu listing all 9 available skills and their lock/enabled status.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jxfsx6612h9ftk82tz9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jxfsx6612h9ftk82tz9.png" alt="List of all skills screen clipping" width="762" height="412"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Skills could also come from installed plugins.&lt;/p&gt;


&lt;h2&gt;
  
  
  8. Code Review via Subagents
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;code review skill&lt;/strong&gt; can delegate each dimension of the review to a dedicated subagent using the Task tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## 1. Security&lt;/span&gt;
Delegate this dimension to the &lt;span class="sb"&gt;`security-analyst`&lt;/span&gt; subagent using the Task tool.

&lt;span class="gu"&gt;## 2. Performance&lt;/span&gt;
Delegate this dimension to the &lt;span class="sb"&gt;`performance-analyst`&lt;/span&gt; subagent using the Task tool.

&lt;span class="gu"&gt;## 3. Test Coverage&lt;/span&gt;
Delegate this dimension to the &lt;span class="sb"&gt;`test-coverage-analyst`&lt;/span&gt; subagent using the Task tool.

&lt;span class="gu"&gt;## 4. Code Quality&lt;/span&gt;
Delegate this dimension to the &lt;span class="sb"&gt;`code-quality-analyst`&lt;/span&gt; subagent using the Task tool.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  How to use a skill
&lt;/h3&gt;

&lt;p&gt;Example prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;❯ look at the auth module and tell me if there are any problems&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The code review skill should auto-invoke:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude auto-invoking the code-review skill and launching four parallel review agents against the auth module.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe9ed10u8gsji0sc5f1s9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe9ed10u8gsji0sc5f1s9.png" alt="Code review skill auto-invoked screen clipping" width="800" height="596"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No explicit integration is needed — it's auto-invoked. No need to use &lt;code&gt;/command&lt;/code&gt; explicitly.&lt;/p&gt;

&lt;p&gt;A sample of the code review report:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Consolidated code review report table listing critical and high-severity findings across the auth module.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6udqpubhuzzc014pz8ca.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6udqpubhuzzc014pz8ca.png" alt="Code review report sample screen clipping 1" width="800" height="582"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Detailed write-up of two critical findings: a syntax-breaking stray await and a bypassed bcrypt password check.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F14alq81al21sbqzux6eg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F14alq81al21sbqzux6eg.png" alt="Code review report sample screen clipping 2" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Example: skill — &lt;code&gt;deploy&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;If I want to deploy, &lt;code&gt;deploy.sh&lt;/code&gt; should be used — so we use a deploy skill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deploy&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Deploy the application to staging. Use when the user says "deploy",&lt;/span&gt;
  &lt;span class="s"&gt;"ship it", "push to staging", or "release". Runs pre-flight checks,&lt;/span&gt;
  &lt;span class="s"&gt;builds, tags, and verifies the deployment. Always confirm with the&lt;/span&gt;
  &lt;span class="s"&gt;user before executing.&lt;/span&gt;
&lt;span class="na"&gt;disable-model-invocation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="gh"&gt;# Deploy Skill&lt;/span&gt;
Deployment logic lives in &lt;span class="sb"&gt;`scripts/deploy.sh`&lt;/span&gt;. That script is the
source of truth — do not re-implement its steps here.

&lt;span class="gu"&gt;## Before Running&lt;/span&gt;
Confirm with the user:
&lt;span class="p"&gt;-&lt;/span&gt; "Ready to deploy to staging. This will tag and push. Proceed?"
&lt;span class="p"&gt;-&lt;/span&gt; Do not proceed without explicit confirmation.

&lt;span class="gu"&gt;## Execution&lt;/span&gt;
Run: &lt;span class="sb"&gt;`bash scripts/deploy.sh`&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Example: skill — &lt;code&gt;write-tests&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Similarly, we can use a write-tests skill to write a test case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;write-tests&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Use this skill whenever new testable logic is added or modified — functions,&lt;/span&gt;
  &lt;span class="s"&gt;methods, classes, route handlers, or modules in src/. Triggers on: "add a&lt;/span&gt;
  &lt;span class="s"&gt;function", "create a utility", "implement this", "write a handler", "add a&lt;/span&gt;
  &lt;span class="s"&gt;route", or any task that produces new logic with inputs and outputs or&lt;/span&gt;
  &lt;span class="s"&gt;side effects. Do NOT wait to be asked — write tests as part of completing&lt;/span&gt;
  &lt;span class="s"&gt;the task. Skip this skill only for: config constants, type definitions,&lt;/span&gt;
  &lt;span class="s"&gt;pure re-exports, or framework boilerplate with no logic.&lt;/span&gt;
&lt;span class="s"&gt;---&lt;/span&gt;
&lt;span class="gh"&gt;# Write Tests Skill&lt;/span&gt;
You are a senior QA engineer writing exhaustive Jest test suites for a
Node.js application.

&lt;span class="gu"&gt;## Stack &amp;amp; Conventions&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Framework**&lt;/span&gt;: Jest with Supertest for HTTP integration tests
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Test location**&lt;/span&gt;: &lt;span class="sb"&gt;`tests/`&lt;/span&gt; mirroring &lt;span class="sb"&gt;`src/`&lt;/span&gt; structure, &lt;span class="sb"&gt;`.test.js`&lt;/span&gt; suffix
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Logger**&lt;/span&gt;: Structured JSON — never assert on &lt;span class="sb"&gt;`console`&lt;/span&gt; output
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Auth**&lt;/span&gt;: &lt;span class="sb"&gt;`JWT_SECRET`&lt;/span&gt; defaults to &lt;span class="sb"&gt;`'dev-secret-key'`&lt;/span&gt; in tests — do not
  assert real security guarantees against this value
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Test data**&lt;/span&gt;: Define inline for simple cases; use &lt;span class="sb"&gt;`tests/fixtures/`&lt;/span&gt; for
  objects reused across 3+ tests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Initiating the write-tests skill:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The write-tests skill identifying and filling a missing leap-year test case for validateDateOfBirth.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrg4tn079eiq9dfkz3wv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrg4tn079eiq9dfkz3wv.png" alt="Write tests skill screen clipping" width="800" height="407"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Skills are auto-invoked. LCOV coverage output:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Test suite results showing 76 passing tests with 100% coverage on validators.js.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfc947ew7k9rmm171owe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfc947ew7k9rmm171owe.png" alt="LCOV coverage screen clipping" width="800" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Creating a custom agent
&lt;/h3&gt;

&lt;p&gt;Skills, agents, and commands differ as follows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent&lt;/strong&gt; — a specialized assistant that does a task in the background and comes up with an answer. Run it separately in the background, or run it via a skill — a skill can also hand a task over to an agent.&lt;/li&gt;
&lt;li&gt;In an agent definition, we can specify exactly &lt;strong&gt;what tools the agent can use&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example of agents:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The project's custom subagent definitions listed in the .claude/agents folder.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyy3p6oq5oz8jsgutfn9b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyy3p6oq5oz8jsgutfn9b.png" alt="Example of agents screen clipping" width="336" height="166"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Example: custom agent — &lt;code&gt;security-analyst&lt;/code&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;security-analyst&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Use this agent for security reviews: authentication bypass, injection vulnerabilities, OWASP Top 10, token handling, and input validation. Invoke with: ask the security-analyst to audit this file.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Bash(npm audit *), Bash(grep *), Write, Edit&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;opus&lt;/span&gt;
&lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;project&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="gh"&gt;# Security Analyst Agent&lt;/span&gt;
You are a senior application security engineer. You read code exclusively through
a security lens. You do not suggest feature improvements or code style changes.
You find vulnerabilities and you explain how to fix them.

&lt;span class="gu"&gt;## Your Focus Areas&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Authentication bypass and broken access control
&lt;span class="p"&gt;-&lt;/span&gt; Injection vulnerabilities: SQL, command, path traversal
&lt;span class="p"&gt;-&lt;/span&gt; Sensitive data exposure in logs, responses, or error messages
&lt;span class="p"&gt;-&lt;/span&gt; Insecure token handling: weak secrets, missing expiry, improper storage
&lt;span class="p"&gt;-&lt;/span&gt; Missing input validation and sanitisation
&lt;span class="p"&gt;-&lt;/span&gt; Dependency vulnerabilities (flag for npm audit review)

&lt;span class="gu"&gt;## OWASP Categorisation&lt;/span&gt;
Tag every finding with its OWASP Top 10 category where applicable.

&lt;span class="gu"&gt;## Memory Protocol&lt;/span&gt;
Before starting, read &lt;span class="sb"&gt;`MEMORY.md`&lt;/span&gt; (an index of one-line links to note files) and
skim any linked notes whose description looks relevant to the file(s) you're
about to audit.

After completing your audit, write to memory only if you found something
non-obvious, recurring, or specific to this codebase that would save real time
on a future audit — not routine findings you'd already report. If so:
&lt;span class="p"&gt;1.&lt;/span&gt; Create a new note file named for the pattern (e.g. &lt;span class="sb"&gt;`jwt-secret-hardcoded-in-config.md`&lt;/span&gt;)
   with &lt;span class="sb"&gt;`name`&lt;/span&gt;, &lt;span class="sb"&gt;`description`&lt;/span&gt;, and &lt;span class="sb"&gt;`metadata: {type: feedback}`&lt;/span&gt; frontmatter, following
   the structure used in existing notes.
&lt;span class="p"&gt;2.&lt;/span&gt; Add a one-line link to it from &lt;span class="sb"&gt;`MEMORY.md`&lt;/span&gt;.

Do not rewrite or duplicate an existing note — update it in place if the same
pattern recurs with new detail.

&lt;span class="gu"&gt;## Output Format&lt;/span&gt;
For each finding:
&lt;span class="p"&gt;-&lt;/span&gt; OWASP Category
&lt;span class="p"&gt;-&lt;/span&gt; Severity: Critical / High / Medium / Low
&lt;span class="p"&gt;-&lt;/span&gt; File and line number
&lt;span class="p"&gt;-&lt;/span&gt; What an attacker could do with this vulnerability
&lt;span class="p"&gt;-&lt;/span&gt; Exact fix with corrected code snippet

End with: total finding count by severity, and one recommended immediate action.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;How does a skill hand over a task to an agent?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude launching the security-analyst subagent in the background to audit the authentication codebase.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbsizi0fu5x22fd4lze1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbbsizi0fu5x22fd4lze1.png" alt="Skill handover to agent screen clipping" width="799" height="286"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The agent runs in the background
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;● Agent "Security audit of auth codebase" finished · 4m 42s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The security-analyst audit found &lt;strong&gt;18 findings&lt;/strong&gt; across &lt;code&gt;authService.js&lt;/code&gt;, &lt;code&gt;tokenHelper.js&lt;/code&gt;, &lt;code&gt;routes.js&lt;/code&gt;, &lt;code&gt;middleware.js&lt;/code&gt;, and &lt;code&gt;validators.js&lt;/code&gt;. Two are critical and one of them also breaks the server outright.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Critical&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Auth bypass&lt;/strong&gt; — &lt;code&gt;bcrypt.compare(...)&lt;/code&gt; at &lt;code&gt;authService.js:32&lt;/code&gt; is called without &lt;code&gt;await&lt;/code&gt;, so the truthy Promise always passes the check. Every password is accepted for any login.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hardcoded JWT fallback secret&lt;/strong&gt; — &lt;code&gt;'dev-secret-key'&lt;/code&gt; in both &lt;code&gt;authService.js:11&lt;/code&gt; and &lt;code&gt;tokenHelper.js:8&lt;/code&gt;, used silently if &lt;code&gt;JWT_SECRET&lt;/code&gt; is unset. Anyone can forge admin tokens.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;High&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Login always builds an inline &lt;code&gt;mockUserRecord&lt;/code&gt; with a constant &lt;code&gt;id: 'user-001'&lt;/code&gt; and attacker-supplied email — every session collapses to the same identity, with a spoofable email claim.&lt;/li&gt;
&lt;li&gt;No rate limiting on &lt;code&gt;/login&lt;/code&gt; or &lt;code&gt;/refresh&lt;/code&gt; — unlimited credential stuffing / refresh-token brute-forcing.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;REFRESH_EXPIRES_IN&lt;/code&gt; is declared but never enforced — refresh tokens never expire.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Medium (6 findings)&lt;/strong&gt; — notably: &lt;code&gt;authService.js:80&lt;/code&gt; is missing the &lt;code&gt;async&lt;/code&gt; keyword on &lt;code&gt;refreshToken&lt;/code&gt;, which is a &lt;code&gt;SyntaxError&lt;/code&gt; that prevents the server from booting at all (confirmed via &lt;code&gt;node --check&lt;/code&gt;) — meaning none of this is currently under test/CI coverage; an inverted expiry comparison in &lt;code&gt;isTokenExpired&lt;/code&gt; that also deletes valid sessions; &lt;code&gt;/logout&lt;/code&gt; can revoke anyone's refresh token (no ownership check); access tokens aren't invalidated on logout; PII (raw emails) logged on every auth attempt; unbounded in-memory token store (memory-exhaustion DoS).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Low (7 findings)&lt;/strong&gt; — untrimmed email reaching downstream sinks, internal error messages leaked to clients via substring matching, missing JWT algorithm allowlist, missing security headers/CORS/body-size limits, loose &lt;code&gt;validateBody&lt;/code&gt; treating &lt;code&gt;null&lt;/code&gt;/&lt;code&gt;0&lt;/code&gt;/&lt;code&gt;false&lt;/code&gt; as "present," a low-severity uuid advisory (not reachable, since &lt;code&gt;uuidv4()&lt;/code&gt; is called with no args), and a latent timing side-channel for user enumeration.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Good news: the newly added &lt;code&gt;validateDateOfBirth&lt;/code&gt; (commit &lt;code&gt;de91f6f&lt;/code&gt;) was checked directly and is clean — correctly rejects rollover dates, future dates, ages over 120, and malformed input. It's not currently wired into any route, so it's not a live attack surface either way.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Recommended immediate fix (one commit):&lt;/strong&gt; add the missing &lt;code&gt;await&lt;/code&gt; on &lt;code&gt;bcrypt.compare&lt;/code&gt; (line 32) and the missing &lt;code&gt;async&lt;/code&gt; on &lt;code&gt;refreshToken&lt;/code&gt; (line 80) together — the second is required just to get the app to boot, and without it nothing else can even be tested. Rotate &lt;code&gt;JWT_SECRET&lt;/code&gt; and make it mandatory at startup right after.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. Multi-Agent Harness Concepts
&lt;/h2&gt;

&lt;p&gt;When Claude Code itself uses multiple subagents to work on a task, we can say it's using a &lt;strong&gt;multi-agent harness&lt;/strong&gt;. The main thread uses the info from all the agents to create a single report.&lt;/p&gt;

&lt;h3&gt;
  
  
  A second pass: performance, test coverage, and code quality
&lt;/h3&gt;

&lt;p&gt;Prompt: &lt;em&gt;"Create agent for performance review, test coverage, and code quality — use the code review skill."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude clarifying review scope, then running performance, test-coverage, and code-quality agents in parallel while fixing the critical bugs.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmjotxcqcpu6cv7l7t99.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmjotxcqcpu6cv7l7t99.png" alt="Performance/test-coverage/code-quality review screen clipping" width="799" height="428"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Full output — Code Review Report (Performance / Test Coverage / Code Quality)&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Context: Re-run of the performance, test-coverage, and code-quality dimensions against &lt;code&gt;src/auth/authService.js&lt;/code&gt; (security was already covered in the prior review). Findings are consistent with the earlier pass; a few additional specifics surfaced this time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Finding&lt;/th&gt;
&lt;th&gt;Severity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Code Quality / Test Coverage&lt;/td&gt;
&lt;td&gt;Missing &lt;code&gt;async&lt;/code&gt; on &lt;code&gt;refreshToken&lt;/code&gt; (line 80) + &lt;code&gt;await&lt;/code&gt; at line 92 → &lt;code&gt;SyntaxError&lt;/code&gt;, module can't load, entire app fails to boot&lt;/td&gt;
&lt;td&gt;🔴 Critical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Code Quality / Test Coverage&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;bcrypt.compare&lt;/code&gt; not awaited (line 32) → auth bypass, confirmed independently by ESLint's &lt;code&gt;require-await&lt;/code&gt; flagging &lt;code&gt;loginUser&lt;/code&gt; as having no real await&lt;/td&gt;
&lt;td&gt;🔴 Critical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Test Coverage&lt;/td&gt;
&lt;td&gt;No test file exists for &lt;code&gt;authService.js&lt;/code&gt; at all — every exported function has 0% coverage&lt;/td&gt;
&lt;td&gt;🔴 Critical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Test Coverage&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;isTokenExpired&lt;/code&gt; inverted logic (line 68) — verified via failing Jest assertions against real signed JWTs&lt;/td&gt;
&lt;td&gt;🔴 Critical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;isTokenExpired&lt;/code&gt; duplicates &lt;code&gt;tokenHelper.js&lt;/code&gt;'s JWT logic instead of delegating (violates CLAUDE.md's documented layering), and is applied to opaque UUID refresh tokens where it always short-circuits via &lt;code&gt;jwt.decode() === null&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;🟠 High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;REFRESH_EXPIRES_IN&lt;/code&gt; and &lt;code&gt;createdAt&lt;/code&gt; are dead code — confirmed via ESLint &lt;code&gt;no-unused-vars&lt;/code&gt; — refresh tokens never actually expire despite CLAUDE.md documenting a 7-day window&lt;/td&gt;
&lt;td&gt;🟠 High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Unbounded &lt;code&gt;refreshTokenStore&lt;/code&gt; Map — no TTL sweep, grows with cumulative logins not concurrent sessions&lt;/td&gt;
&lt;td&gt;🟠 High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Test Coverage&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;refreshToken()&lt;/code&gt;, &lt;code&gt;revokeToken()&lt;/code&gt;, &lt;code&gt;generateRefreshToken()&lt;/code&gt;, &lt;code&gt;generateAccessToken()&lt;/code&gt; have no direct unit tests or edge-case coverage&lt;/td&gt;
&lt;td&gt;🟠 High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;bcryptjs&lt;/code&gt; (pure-JS, not native bcrypt) runs password comparison on the main thread, not the libuv threadpool — inflates tail latency under concurrent login bursts&lt;/td&gt;
&lt;td&gt;🟡 Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;Swallowed error in &lt;code&gt;isTokenExpired&lt;/code&gt;'s &lt;code&gt;catch (err)&lt;/code&gt; (line 69-71) — caught but never logged, inconsistent with the rest of the file's logging pattern&lt;/td&gt;
&lt;td&gt;🟡 Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;curly&lt;/code&gt; ESLint rule violation on line 64 (&lt;code&gt;if (...) return true;&lt;/code&gt; without braces)&lt;/td&gt;
&lt;td&gt;🟡 Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;require-await&lt;/code&gt; violation on &lt;code&gt;getUserById&lt;/code&gt; (line 153) — declared async with no await inside&lt;/td&gt;
&lt;td&gt;🟡 Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;Test Coverage&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;loginUser&lt;/code&gt; with no &lt;code&gt;userRecord&lt;/code&gt;, JWT claims shape, &lt;code&gt;getUserById&lt;/code&gt; stub contract — untested but low-risk edge cases&lt;/td&gt;
&lt;td&gt;🟢 Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;revokeToken&lt;/code&gt;'s "not found" branch logs nothing, inconsistent with other failure branches in the file&lt;/td&gt;
&lt;td&gt;🟢 Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;Code Quality&lt;/td&gt;
&lt;td&gt;Imprecise JSDoc return shapes (&lt;code&gt;@returns {object}&lt;/code&gt; instead of a typed shape)&lt;/td&gt;
&lt;td&gt;🟢 Low&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;New since the last pass&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tooling is broken at the project level: &lt;code&gt;npm run lint&lt;/code&gt; currently fails outright — the repo ships &lt;code&gt;.eslintrc.json&lt;/code&gt; (legacy config) but has ESLint 10.9.1 installed, which requires flat &lt;code&gt;eslint.config.*&lt;/code&gt;. No flat config exists at the repo root. The code-quality agent worked around this by building an equivalent flat config in the scratchpad to get real lint signal (never touched the actual repo files), but as things stand today, &lt;code&gt;npm run lint&lt;/code&gt; gives false confidence — it fails before it even reaches your code. This is worth fixing independent of &lt;code&gt;authService.js&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;bcryptjs&lt;/code&gt; vs native bcrypt (finding 9): not raised as a distinct finding in the first pass — flagged now because it compounds with the login-bypass fix. Once &lt;code&gt;await&lt;/code&gt; is added back (finding 2), every login request will run a real CPU-bound bcrypt comparison on the main thread; under concurrent login bursts this inflates latency for unrelated concurrent requests. Not urgent, but worth knowing before assuming the fix is "free."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Summary Score: 2/10 (unchanged)&lt;/strong&gt; — the file can't boot, has zero test coverage, and now additionally the project's lint tooling can't even verify it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommended First Action:&lt;/strong&gt; unchanged — fix the missing &lt;code&gt;async&lt;/code&gt; on &lt;code&gt;refreshToken&lt;/code&gt; (line 80) together with the missing &lt;code&gt;await&lt;/code&gt; on &lt;code&gt;bcrypt.compare&lt;/code&gt; (line 32) in one commit — nothing else here (tests, lint, the rest of the app) is reachable until the module can actually load. Separately, worth a quick fix to add &lt;code&gt;eslint.config.js&lt;/code&gt; so &lt;code&gt;npm run lint&lt;/code&gt; works again — that's a repo-wide gap, not specific to this file.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is a &lt;strong&gt;hub and spoke model&lt;/strong&gt; — using subagents via a skill. However, what we actually want in some cases is &lt;strong&gt;agents communicating with each other independently in the background&lt;/strong&gt; to finish a task; the hub-and-spoke approach above cannot do that. Agents can instead be run using other agentic harnesses — loop, schedule, etc.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Difference between agent and agentic
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Comparison diagram contrasting framework-driven "Agent" orchestration with model-native "Agentic" execution.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7quqv62a70qls0bn9nt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7quqv62a70qls0bn9nt.png" alt="Agent vs agentic diagram screen clipping" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Goal&lt;/strong&gt;, &lt;strong&gt;loop&lt;/strong&gt;, &lt;strong&gt;schedule&lt;/strong&gt;, &lt;strong&gt;workflow&lt;/strong&gt;, &lt;strong&gt;batch&lt;/strong&gt; — agentic primitives available which help create subtasks (e.g. ultraplan, autofix-pr).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/commands&lt;/code&gt; are manual.&lt;/li&gt;
&lt;li&gt;Skills get applied automatically based on the prompt.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/agent&lt;/code&gt; runs in the background.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  10. Hooks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;HOOKS, PLUGINS, MCP e2e flow&lt;/strong&gt; — from evaluating code to fixing code, full auto flow.&lt;/p&gt;

&lt;p&gt;Lifecycle methods available in Claude: from the start to the end of a session, internal events fire inside the session, which are used as plug points for our own code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hook event flow
&lt;/h3&gt;

&lt;p&gt;Whenever Claude reaches a lifecycle boundary — e.g. when a tool is about to run — it spawns our custom script as a subprocess and feeds it a JSON blob on stdin. The script can execute anything, then return:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an exit code (&lt;code&gt;0&lt;/code&gt; or &lt;code&gt;2&lt;/code&gt;), or&lt;/li&gt;
&lt;li&gt;a JSON object on stdout containing a structured decision, reason, and context for Claude to read.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hooks can be created using Python or JavaScript. They take input from Claude Code, display it on screen, and allow the tool to run.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example: &lt;code&gt;PreToolUse&lt;/code&gt; hook on Bash
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PreToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node .claude/hooks/bash-guard.js"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node .claude/hooks/pretooluse-demo.js"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Since this hook sits directly in front of every Bash call, it's a good interception point for anything you want enforced before a shell command runs. Some concrete uses, especially relevant to a repo like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails / safety&lt;/strong&gt; (this is what &lt;code&gt;bash-guard.js&lt;/code&gt; already seems to do) — block dangerous patterns beyond what &lt;code&gt;settings.json&lt;/code&gt;'s deny-list covers, e.g. &lt;code&gt;rm -rf&lt;/code&gt; variants, &lt;code&gt;git push --force&lt;/code&gt;, piping to &lt;code&gt;sh&lt;/code&gt;, etc., using regex instead of exact string matching.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  All hook handler types
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PreToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node .claude/hooks/pretooluse-demo.js"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://www.example.com/validate"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Evaluate the steps before running the tool"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Run security-analyst to verify the issues"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mcp_tool"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"server"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"some-mcp-server"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"validate_edit"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The bash-guard hook's log file recording every Bash command it allowed during the session.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frw6yxiutl9hbbhxy54ux.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frw6yxiutl9hbbhxy54ux.png" alt="Hook handler types screen clipping" width="800" height="186"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Blocking dangerous commands
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DANGER_PATTERNS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sr"&gt;/rm&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+-rf&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                          &lt;span class="c1"&gt;// recursive force-delete from root (e.g. "rm -rf /")&lt;/span&gt;
    &lt;span class="sr"&gt;/sudo&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+rm/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                              &lt;span class="c1"&gt;// any sudo-elevated delete — bypasses normal permission checks&lt;/span&gt;
    &lt;span class="sr"&gt;/:&lt;/span&gt;&lt;span class="se"&gt;\(\)\{\s&lt;/span&gt;&lt;span class="sr"&gt;*:&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="sr"&gt;:&amp;amp;&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;\}\s&lt;/span&gt;&lt;span class="sr"&gt;*;&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;*:/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;           &lt;span class="c1"&gt;// fork bomb — ":(){ :|:&amp;amp; };:" spawns processes until the system locks up&lt;/span&gt;
    &lt;span class="sr"&gt;/dd&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+if=.*of=&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;dev&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;sd/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                 &lt;span class="c1"&gt;// raw disk write via dd — can overwrite an entire drive&lt;/span&gt;
    &lt;span class="sr"&gt;/chmod&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+777&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;// world-writable permissions on root — opens up the whole filesystem&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Notification on Slack
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"Notification"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"permission_prompt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node .claude/hooks/slack-notify.js"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Logs can be used to audit our work — when we did what, and how much time we spent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sharing configuration
&lt;/h3&gt;

&lt;p&gt;What if I want to use all my settings in the &lt;code&gt;.claude&lt;/code&gt; folder in another project? This can be done via GitHub. But how can we give it to a third party? Through &lt;strong&gt;plugins&lt;/strong&gt; — see &lt;a href="https://code.claude.com/docs/en/plugins" rel="noopener noreferrer"&gt;code.claude.com/docs/en/plugins&lt;/a&gt; for how to create one.&lt;/p&gt;




&lt;h2&gt;
  
  
  11. MCP (Model Context Protocol)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The concept of an MCP server
&lt;/h3&gt;

&lt;p&gt;The GitHub REST API knows how to connect to GitHub on your behalf.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Diagram showing Git Bash and the GH CLI both reaching GitHub's SaaS platform through the GitHub REST API.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccqi23ew433pu5lby7w3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccqi23ew433pu5lby7w3.png" alt="GitHub REST API screen clipping" width="502" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Either Claude directly connects via Bash, or via the &lt;code&gt;gh&lt;/code&gt; CLI — both of which need to be authenticated. Actually, the &lt;code&gt;gh&lt;/code&gt; CLI is itself using Bash commands under the hood. But can Claude Code directly connect to the GitHub REST API to be faster and more efficient? &lt;code&gt;curl&lt;/code&gt; can be used for that:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Expanded diagram showing how Claude Code uses its Bash tool (via curl, git, or gh CLI) to reach the GitHub REST API.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3c6l5i5p3myzgsw918c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3c6l5i5p3myzgsw918c.png" alt="Curl-based GH REST API connection screen clipping" width="800" height="409"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why MCP exists
&lt;/h3&gt;

&lt;p&gt;Every ecosystem/SaaS (like GitHub) has a separate, specific, proprietary way of connecting to it — so if Claude has to connect to multiple SaaS platforms, it would need to know all of them, which deviates Claude from its core functionality.&lt;/p&gt;

&lt;p&gt;Can GitHub give a way to connect directly to a REST API? So Anthropic created a new protocol:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Claude Code has an &lt;strong&gt;inbuilt MCP client&lt;/strong&gt; when installed.&lt;/li&gt;
&lt;li&gt;GitHub/any SaaS provider provides an &lt;strong&gt;MCP server&lt;/strong&gt; — basically an endpoint.&lt;/li&gt;
&lt;li&gt;Our job is only to connect to the MCP server; the MCP server handles the rest.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Three types of MCP server
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;stdio&lt;/strong&gt; — runs locally on the machine&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;http&lt;/strong&gt; — hosted at an HTTP location&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SSE&lt;/strong&gt; — Server-Sent Events, a streaming protocol&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;All AI tools support MCP servers. The MCP client is part of the ecosystem.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Claude Code runs with an agentic harness — the LLM decides what to do, and the LLM tells Claude Code which tools to use.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  12. End-to-End Automation
&lt;/h2&gt;

&lt;p&gt;End-to-end automation of security review → issue creation → issue fix → PR creation → PR fix, all through agents.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The /security-review command loading its skill to audit tokenHelper.js.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsecislasgjmafpge71ah.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsecislasgjmafpge71ah.png" alt="End-to-end automation intro screen clipping" width="584" height="93"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Security review finding a critical hardcoded fallback JWT signing secret vulnerability in tokenHelper.js.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoo4jschokn8ljjgfryl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoo4jschokn8ljjgfryl.png" alt="Automation flow screen clipping" width="800" height="174"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;/create-issue&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Claude using the gh CLI to file a GitHub issue for the hardcoded JWT secret vulnerability.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flw19l10sb7c8eyw6k7l1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flw19l10sb7c8eyw6k7l1.png" alt="Create issue screen clipping" width="800" height="304"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;GitHub issue #2 detailing the hardcoded fallback JWT secret bug, its impact, and expected fix.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg3iy5j7gpzpo5cze2hi4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg3iy5j7gpzpo5cze2hi4.png" alt="Create issue result screen clipping" width="800" height="682"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Now fix the issue
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Claude implementing the fix: removing the hardcoded JWT secret fallback and requiring the env var explicitly.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr2aaxl64ntvxmah7wgtg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr2aaxl64ntvxmah7wgtg.png" alt="Fix the issue screen clipping" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Commit the changes
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;The /commit command staging, reviewing, and pushing the JWT secret fix with a detailed commit message.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foxh3fk45tvwrtqkx79vp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foxh3fk45tvwrtqkx79vp.png" alt="Commit changes screen clipping" width="800" height="532"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;GitHub activity showing the fix commit automatically linked to close issue #2.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxq6ggwpjdlmjbxqocre.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxq6ggwpjdlmjbxqocre.png" alt="Commit result screen clipping 1" width="798" height="171"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Pull request description summarizing what changed and why for the JWT secret fix.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs8sdrtyg52zwzg139t09.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs8sdrtyg52zwzg139t09.png" alt="Commit result screen clipping 2" width="800" height="701"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  13. Agent Teams
&lt;/h2&gt;

&lt;p&gt;Beyond single subagent delegation, Claude Code supports &lt;strong&gt;Agent Teams&lt;/strong&gt; — multiple agents working the same task from different angles in parallel, then synthesizing a combined result.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example prompt
&lt;/h3&gt;

&lt;blockquote&gt;
&lt;p&gt;Create an agent team to review PR #1.&lt;br&gt;
Spawn three reviewers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reviewer 1: focused exclusively on security implications&lt;/li&gt;
&lt;li&gt;Reviewer 2: focused on performance and scalability&lt;/li&gt;
&lt;li&gt;Reviewer 3: validating test coverage and edge cases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Have them each review independently, then synthesize findings into a single report.&lt;br&gt;
Create a new branch with custom name including today's date.&lt;br&gt;
Merge and close the PR with comment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;User prompt instructing Claude to spawn three specialized reviewer agents to review and merge PR #1.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3y6axvgkb9davveadh4y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3y6axvgkb9davveadh4y.png" alt="Agent teams prompt screen clipping" width="799" height="209"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The 3 agents are invoked in the background:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude launching three specialized reviewer agents in parallel to review PR #1.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtyyqypghqmxtt8q83t3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtyyqypghqmxtt8q83t3.png" alt="Agents invoked in background screen clipping" width="800" height="392"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Progress log showing the security and test-coverage reviewer agents finishing with real findings.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa61rg7s3odmebbkaz7a8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa61rg7s3odmebbkaz7a8.png" alt="Agent team progress screen clipping 1" width="800" height="465"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude synthesizing all three reviewer agents' findings before applying fixes and merging.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbt0r6fynkc7sj88vjwu9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbt0r6fynkc7sj88vjwu9.png" alt="Agent team progress screen clipping 2" width="800" height="110"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Claude merging the verified fix, filing two follow-up issues, and preparing to post the review synthesis as a PR comment.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjg8b2d7k3m8twcxgyzj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjg8b2d7k3m8twcxgyzj.png" alt="Agent team result screen clipping 1" width="800" height="431"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Final summary of the agent-team review: verdicts from all three reviewers and the actions taken to merge PR #1.&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwep2tfayik2sa7gce8e6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwep2tfayik2sa7gce8e6.png" alt="Agent team result screen clipping 2" width="800" height="390"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  14. Key Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The &lt;strong&gt;agentic loop&lt;/strong&gt; (gather context → decide → act → verify) is the mental model for everything Claude Code does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLAUDE.md&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;.claude/settings.json&lt;/code&gt;&lt;/strong&gt; are how you shape Claude's behavior and permissions for a specific project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commands&lt;/strong&gt; are manual, &lt;strong&gt;Skills&lt;/strong&gt; auto-trigger, and &lt;strong&gt;Agents&lt;/strong&gt; run scoped, background work — often orchestrated together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hooks&lt;/strong&gt; and &lt;strong&gt;MCP&lt;/strong&gt; are the two extension points: hooks intercept lifecycle events for guardrails/automation, MCP standardizes how Claude talks to external tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session management&lt;/strong&gt; (&lt;code&gt;/resume&lt;/code&gt;, &lt;code&gt;/fork&lt;/code&gt;, &lt;code&gt;/compact&lt;/code&gt;) is essential for working on multiple issues without polluting context or wasting tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent Teams&lt;/strong&gt; unlock genuinely parallel, multi-perspective work — like having three reviewers look at a PR simultaneously.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This was a genuinely comprehensive look at how far the agentic coding model has come — from a single loop reading and writing files, to fully orchestrated, permissioned, multi-agent development workflows.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
    <item>
      <title>AI &amp; ML Foundations: Cleaning Data, Building Regression &amp; Classification Models in Python.</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Mon, 31 Aug 2026 16:53:20 +0000</pubDate>
      <link>https://dev.to/yashnigam/ai-ml-foundations-cleaning-data-building-regression-classification-models-in-python-7gc</link>
      <guid>https://dev.to/yashnigam/ai-ml-foundations-cleaning-data-building-regression-classification-models-in-python-7gc</guid>
      <description>&lt;h1&gt;
  
  
  My 5-Day Journey into Machine Learning: From Data Cleaning to My First ML Models
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Learning goal:&lt;/strong&gt; This is not a collection of code snippets to memorize. It is a beginner-friendly guide to understanding &lt;em&gt;why&lt;/em&gt; each step in a typical machine-learning workflow exists, what problem it solves, and how the pieces fit together.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A note about the source material
&lt;/h2&gt;

&lt;p&gt;This guide was built primarily from the four Jupyter notebooks you provided:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;11_Aug_mpg.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;12-Aug-2.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;12th-Aug-Telco-Customer-churn.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;13-Aug.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;123&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your PDF notes were also provided, but the PDF is image-based and its text could not be reliably extracted. So I have &lt;strong&gt;not invented or silently reconstructed&lt;/strong&gt; material from the PDF. The explanations below are grounded mainly in the notebooks, with general ML knowledge added to explain the concepts behind the code.&lt;/p&gt;




&lt;h1&gt;
  
  
  1. Overview: What did I actually learn in these five days?
&lt;/h1&gt;

&lt;p&gt;At first glance, machine learning can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load dataset
    ↓
Clean data
    ↓
Create graphs
    ↓
Train model
    ↓
Check accuracy
    ↓
Done!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But that misses the most important part.&lt;/p&gt;

&lt;p&gt;The real ML workflow is closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Business / real-world question
        ↓
Define what we want to predict
        ↓
Understand the data
        ↓
Clean and prepare the data
        ↓
Explore the data visually
        ↓
Choose useful features
        ↓
Convert data into numbers
        ↓
Split into training and testing data
        ↓
Choose an appropriate ML algorithm
        ↓
Train the model
        ↓
Evaluate it on unseen data
        ↓
Compare models
        ↓
Improve preprocessing / model settings
        ↓
Use the model to make predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important insight is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Machine learning is not mainly about choosing an algorithm. It is about turning a real-world question and messy data into a reliable prediction process.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Your notebooks actually demonstrate most of this workflow.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. What is Artificial Intelligence?
&lt;/h1&gt;

&lt;p&gt;Artificial Intelligence (AI) is the broad idea of building systems that can perform tasks that normally require some form of human intelligence.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;recognizing an image&lt;/li&gt;
&lt;li&gt;understanding language&lt;/li&gt;
&lt;li&gt;recommending a movie&lt;/li&gt;
&lt;li&gt;detecting fraud&lt;/li&gt;
&lt;li&gt;predicting whether a customer will leave&lt;/li&gt;
&lt;li&gt;predicting the price/value of something&lt;/li&gt;
&lt;li&gt;generating text or images&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A useful mental model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Artificial Intelligence
│
├── Machine Learning
│   │
│   ├── Supervised Learning
│   ├── Unsupervised Learning
│   └── Reinforcement Learning
│
├── Deep Learning
│
└── Generative AI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Fyash-nigam%2FAI-ML-Foundations%2Fblob%2F7516a4b681369d4b105ca19706f7d06da65e30a9%2Fimages%2FAIMLHierarchy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Fyash-nigam%2FAI-ML-Foundations%2Fblob%2F7516a4b681369d4b105ca19706f7d06da65e30a9%2Fimages%2FAIMLHierarchy.png" alt="image" width="" height=""&gt;&lt;/a&gt;&lt;br&gt;
This is simplified because these areas overlap, but it is a useful beginner mental model.&lt;/p&gt;


&lt;h1&gt;
  
  
  3. What is Machine Learning?
&lt;/h1&gt;

&lt;p&gt;Machine learning is a way of building systems where the computer learns patterns from data instead of us manually writing every rule.&lt;/p&gt;

&lt;p&gt;Imagine we want to predict whether a telecom customer will churn.&lt;/p&gt;

&lt;p&gt;A traditional rule-based approach might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF customer has a month-to-month contract
AND monthly charges are high
AND tenure is low
THEN churn = yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem is that we have to invent those rules ourselves.&lt;/p&gt;

&lt;p&gt;With machine learning, we give the algorithm historical examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer information → Actual churn result

Customer A → No
Customer B → Yes
Customer C → No
Customer D → Yes
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The algorithm tries to learn the relationship between the input information and the known outcome.&lt;/p&gt;

&lt;p&gt;Then we can give it a new customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New customer information
        ↓
      Model
        ↓
Predicted churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the core idea of supervised machine learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. The most important ML vocabulary
&lt;/h1&gt;

&lt;p&gt;Before learning algorithms, understand these words.&lt;/p&gt;

&lt;h2&gt;
  
  
  4.1 Dataset
&lt;/h2&gt;

&lt;p&gt;A dataset is a collection of examples.&lt;/p&gt;

&lt;p&gt;In your notebooks, a dataset is loaded into a Pandas DataFrame.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Telco_Customer_Churn.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Think of the DataFrame as a spreadsheet.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.2 Row / observation / sample
&lt;/h2&gt;

&lt;p&gt;Each row represents one example.&lt;/p&gt;

&lt;p&gt;For the Telco dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one customer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the Auto MPG dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one car
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the insurance dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one customer/policy record
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These words are often used interchangeably:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;row&lt;/li&gt;
&lt;li&gt;observation&lt;/li&gt;
&lt;li&gt;sample&lt;/li&gt;
&lt;li&gt;instance&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4.3 Feature
&lt;/h2&gt;

&lt;p&gt;A feature is an input variable that the model can use to make a prediction.&lt;/p&gt;

&lt;p&gt;For customer churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tenure
MonthlyCharges
Contract
InternetService
PaymentMethod
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are features.&lt;/p&gt;

&lt;p&gt;A useful question to ask is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What information is available to the model before it makes its prediction?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That information is generally your feature set.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.4 Target
&lt;/h2&gt;

&lt;p&gt;The target is what we are trying to predict.&lt;/p&gt;

&lt;p&gt;Examples from your notebooks:&lt;/p&gt;

&lt;h3&gt;
  
  
  Telco
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Target = Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model predicts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0 → No churn
1 → Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Insurance
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Target = Customer Lifetime Value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model predicts a numerical value.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.5 X and y
&lt;/h2&gt;

&lt;p&gt;You repeatedly used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the most important patterns in supervised ML.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X = inputs/features
Y = target/output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X
├── tenure
├── MonthlyCharges
├── Contract
├── InternetService
└── ...

Y
└── Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;X = what the model knows&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Y = what we want the model to learn to predict&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  5. The three major types of Machine Learning
&lt;/h1&gt;

&lt;h2&gt;
  
  
  5.1 Supervised Learning
&lt;/h2&gt;

&lt;p&gt;In supervised learning, we have examples where the correct answer is already known.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input X → Known answer Y
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model learns the relationship between X and Y.&lt;/p&gt;

&lt;p&gt;Your notebooks mainly use supervised learning.&lt;/p&gt;

&lt;p&gt;There are two major supervised learning tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Classification
&lt;/h3&gt;

&lt;p&gt;The target is a category.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Spam / Not Spam
Churn / No Churn
Fraud / Not Fraud
Disease / No Disease
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your Telco project is a classification problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Regression
&lt;/h3&gt;

&lt;p&gt;The target is a numerical value.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;House price
Temperature
Sales
Customer Lifetime Value
Fuel efficiency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your Customer Lifetime Value project is a regression problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. Classification vs Regression
&lt;/h1&gt;

&lt;p&gt;This distinction is extremely important.&lt;/p&gt;

&lt;h2&gt;
  
  
  Classification
&lt;/h2&gt;

&lt;p&gt;Question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which category does this example belong to?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer → Churn or No Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Possible output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0
1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common classification algorithms you tried:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Logistic Regression&lt;/li&gt;
&lt;li&gt;Support Vector Classifier (SVC)&lt;/li&gt;
&lt;li&gt;Decision Tree Classifier&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Regression
&lt;/h2&gt;

&lt;p&gt;Question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What numerical value should we predict?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer → Customer Lifetime Value = 6,543.21
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common regression algorithms you tried:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Linear Regression&lt;/li&gt;
&lt;li&gt;KNN Regressor&lt;/li&gt;
&lt;li&gt;SVR&lt;/li&gt;
&lt;li&gt;Decision Tree Regressor&lt;/li&gt;
&lt;li&gt;Bagging Regressor&lt;/li&gt;
&lt;li&gt;AdaBoost Regressor&lt;/li&gt;
&lt;li&gt;Random Forest Regressor&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  7. Unsupervised Learning
&lt;/h1&gt;

&lt;p&gt;In unsupervised learning, we do not have a known target.&lt;/p&gt;

&lt;p&gt;Instead, the algorithm tries to discover structure in the data.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer data
     ↓
Find groups
     ↓
Group 1: high-value customers
Group 2: price-sensitive customers
Group 3: new customers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;K-Means clustering&lt;/li&gt;
&lt;li&gt;Hierarchical clustering&lt;/li&gt;
&lt;li&gt;PCA&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You did not build an unsupervised-learning project in these notebooks, but it is important to know where it fits.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. Reinforcement Learning
&lt;/h1&gt;

&lt;p&gt;Reinforcement learning is different.&lt;/p&gt;

&lt;p&gt;An agent interacts with an environment and receives rewards or penalties.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓ action
Environment
  ↓
Reward / penalty
  ↓
Agent learns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;game-playing agents&lt;/li&gt;
&lt;li&gt;robotics&lt;/li&gt;
&lt;li&gt;certain recommendation/control systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This was not part of your five-day practical work.&lt;/p&gt;




&lt;h1&gt;
  
  
  9. Different ways to approach an ML problem
&lt;/h1&gt;

&lt;p&gt;A useful way to think about ML is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which algorithm should I use?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead ask:&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1 — What is the question?
&lt;/h3&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I predict whether a telecom customer will churn?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Step 2 — What type of target do I have?
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn → category → Classification
Customer Lifetime Value → number → Regression
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 3 — What data is available?
&lt;/h3&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What columns exist?&lt;/li&gt;
&lt;li&gt;Which are numerical?&lt;/li&gt;
&lt;li&gt;Which are categorical?&lt;/li&gt;
&lt;li&gt;Are values missing?&lt;/li&gt;
&lt;li&gt;Are there strange values?&lt;/li&gt;
&lt;li&gt;Are some columns identifiers?&lt;/li&gt;
&lt;li&gt;Is the target balanced?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 4 — Explore before modeling
&lt;/h3&gt;

&lt;p&gt;Use statistics and graphs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5 — Prepare the data
&lt;/h3&gt;

&lt;p&gt;Convert it into a form the algorithm can understand.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6 — Establish a baseline
&lt;/h3&gt;

&lt;p&gt;Train a simple model first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 7 — Compare alternatives
&lt;/h3&gt;

&lt;p&gt;Try other appropriate models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 8 — Evaluate on unseen data
&lt;/h3&gt;

&lt;p&gt;A model is useful only if it works beyond the data it memorized.&lt;/p&gt;




&lt;h1&gt;
  
  
  10. Google Colab: Why did we use it?
&lt;/h1&gt;

&lt;p&gt;Google Colab gives you a cloud-based Jupyter Notebook environment.&lt;/p&gt;

&lt;p&gt;Instead of installing everything locally, you can run Python code in the browser.&lt;/p&gt;

&lt;p&gt;A notebook combines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Explanation
+
Python code
+
Output
+
Graphs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes it very useful for learning ML.&lt;/p&gt;

&lt;p&gt;Your notebooks use paths such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;auto&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;mpg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;csv&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is typical of a Colab environment.&lt;/p&gt;

&lt;p&gt;A normal learning workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Upload dataset
      ↓
Open notebook
      ↓
Run cells
      ↓
Inspect output
      ↓
Change code
      ↓
Run again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  11. The Python libraries you used
&lt;/h1&gt;

&lt;h2&gt;
  
  
  11.1 NumPy
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NumPy provides numerical operations and arrays.&lt;/p&gt;

&lt;p&gt;You used it for things such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;
&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expm1&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;code&gt;np.nan&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Represents a missing numerical value.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;np.log1p(x)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Calculates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;log(1 + x)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is useful when a target is heavily skewed.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;np.expm1(x)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Reverses &lt;code&gt;log1p&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exp(x) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You used this later to convert log predictions back to the original CLV scale.&lt;/p&gt;




&lt;h1&gt;
  
  
  12. Pandas
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pandas is one of the most important Python libraries for data work.&lt;/p&gt;

&lt;p&gt;Think of a Pandas DataFrame as a spreadsheet that Python can manipulate.&lt;/p&gt;

&lt;p&gt;You used it to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;load CSV files&lt;/li&gt;
&lt;li&gt;inspect data&lt;/li&gt;
&lt;li&gt;remove columns&lt;/li&gt;
&lt;li&gt;convert data types&lt;/li&gt;
&lt;li&gt;detect missing values&lt;/li&gt;
&lt;li&gt;fill missing values&lt;/li&gt;
&lt;li&gt;select X and Y&lt;/li&gt;
&lt;li&gt;create encoded columns&lt;/li&gt;
&lt;li&gt;save predictions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Telco_Customer_Churn.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  13. Matplotlib
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;matplotlib.pyplot&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;plt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Matplotlib is a general-purpose plotting library.&lt;/p&gt;

&lt;p&gt;It provides the foundation for many Python visualizations.&lt;/p&gt;

&lt;p&gt;You used it for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;figure&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;title&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;show&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  14. Seaborn
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;seaborn&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sns&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seaborn makes statistical visualizations easier to create.&lt;/p&gt;

&lt;p&gt;You used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;countplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;boxplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;histplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pairplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;scatterplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;heatmap&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A key lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A graph is not decoration. It is a tool for asking questions about the data.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  15. Scikit-learn
&lt;/h1&gt;

&lt;p&gt;Scikit-learn is the main ML library used in your notebooks.&lt;/p&gt;

&lt;p&gt;You used it for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;splitting data&lt;/li&gt;
&lt;li&gt;preprocessing&lt;/li&gt;
&lt;li&gt;encoding&lt;/li&gt;
&lt;li&gt;scaling&lt;/li&gt;
&lt;li&gt;classification&lt;/li&gt;
&lt;li&gt;regression&lt;/li&gt;
&lt;li&gt;evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The basic pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SomeModel&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;predictions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This pattern appears again and again in ML.&lt;/p&gt;




&lt;h1&gt;
  
  
  16. The most important workflow: EDA
&lt;/h1&gt;

&lt;p&gt;EDA means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Exploratory Data Analysis&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;EDA is the process of investigating the dataset before building the final model.&lt;/p&gt;

&lt;p&gt;Your notebooks use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;plus graphs.&lt;/p&gt;

&lt;p&gt;These are not random commands.&lt;/p&gt;

&lt;p&gt;They answer different questions.&lt;/p&gt;




&lt;h1&gt;
  
  
  17. &lt;code&gt;df.head()&lt;/code&gt; — What does the data look like?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It displays the first few rows.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because before doing anything else, you want to see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;column names&lt;/li&gt;
&lt;li&gt;example values&lt;/li&gt;
&lt;li&gt;obvious data problems&lt;/li&gt;
&lt;li&gt;whether the data loaded correctly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Think of it as opening the box before using what is inside.&lt;/p&gt;




&lt;h1&gt;
  
  
  18. &lt;code&gt;df.shape&lt;/code&gt; — How much data do I have?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Auto MPG:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(398, 9)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;398 rows
9 columns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why does this matter?&lt;/p&gt;

&lt;p&gt;Because dataset size affects:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model choice&lt;/li&gt;
&lt;li&gt;computation time&lt;/li&gt;
&lt;li&gt;confidence in results&lt;/li&gt;
&lt;li&gt;risk of overfitting&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  19. &lt;code&gt;df.info()&lt;/code&gt; — What types of data do I have?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This tells you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;column names&lt;/li&gt;
&lt;li&gt;number of non-null values&lt;/li&gt;
&lt;li&gt;data types&lt;/li&gt;
&lt;li&gt;memory usage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is extremely important.&lt;/p&gt;

&lt;p&gt;For example, in Auto MPG:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;horsepower → object
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance horsepower sounds numerical.&lt;/p&gt;

&lt;p&gt;But Pandas sees it as text.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because the dataset contains values such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A column containing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;130
165
150
?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;cannot be treated as a clean numerical column.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;df.info()&lt;/code&gt; helps you detect data-type problems before modeling.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  20. &lt;code&gt;df.describe()&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and sometimes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives summary statistics.&lt;/p&gt;

&lt;p&gt;For numerical data you get things such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;count&lt;/li&gt;
&lt;li&gt;mean&lt;/li&gt;
&lt;li&gt;standard deviation&lt;/li&gt;
&lt;li&gt;minimum&lt;/li&gt;
&lt;li&gt;25th percentile&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;75th percentile&lt;/li&gt;
&lt;li&gt;maximum&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This helps you understand the scale and distribution of variables.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;minimum
   ↓
25%
   ↓
median
   ↓
75%
   ↓
maximum
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A large difference between the 75th percentile and maximum can be a clue that there may be extreme values.&lt;/p&gt;




&lt;h1&gt;
  
  
  21. &lt;code&gt;df.sample()&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or similar.&lt;/p&gt;

&lt;p&gt;This shows random rows.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because the first rows may not represent the whole dataset.&lt;/p&gt;

&lt;p&gt;Random samples can expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unusual values&lt;/li&gt;
&lt;li&gt;formatting issues&lt;/li&gt;
&lt;li&gt;unexpected categories&lt;/li&gt;
&lt;li&gt;data-entry problems&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  22. Missing data
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many missing values are there in each column?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Missing data is important because many ML algorithms cannot directly work with missing values.&lt;/p&gt;

&lt;p&gt;But there is a subtle lesson from your Auto MPG notebook.&lt;/p&gt;

&lt;p&gt;You first checked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while &lt;code&gt;horsepower&lt;/code&gt; still contained &lt;code&gt;"?"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Pandas did not count &lt;code&gt;"?"&lt;/code&gt; as a missing value because &lt;code&gt;"?"&lt;/code&gt; is a string, not a true &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"?" ≠ NaN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So data cleaning sometimes requires identifying &lt;strong&gt;fake missing values&lt;/strong&gt; first.&lt;/p&gt;




&lt;h1&gt;
  
  
  23. Auto MPG: cleaning horsepower
&lt;/h1&gt;

&lt;p&gt;Your notebook did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This converts the placeholder into a real missing value.&lt;/p&gt;

&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now Pandas can treat horsepower as numerical data.&lt;/p&gt;

&lt;p&gt;This is a very important data-cleaning pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;messy text
   ↓
identify invalid placeholder
   ↓
convert to NaN
   ↓
convert column to numeric
   ↓
handle missing values
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  24. Why use the median?
&lt;/h1&gt;

&lt;p&gt;You calculated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;median_horsepower&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and then filled missing values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;median_horsepower&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why median?&lt;/p&gt;

&lt;p&gt;Because the median is less affected by extreme values than the mean.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10, 11, 12, 13, 1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mean is pulled heavily upward by 1000.&lt;/p&gt;

&lt;p&gt;The median is much more representative of the middle of the data.&lt;/p&gt;

&lt;p&gt;A useful rule of thumb:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fairly symmetric data → mean may work well&lt;/li&gt;
&lt;li&gt;skewed data / outliers → median is often safer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;There is no universal rule that "median is always correct."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The best treatment depends on the data and the problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  25. Dropping an identifier
&lt;/h1&gt;

&lt;p&gt;You considered dropping:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;car name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and actually dropped &lt;code&gt;customerID&lt;/code&gt; in the Telco work.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;An identifier is usually not a meaningful predictive feature.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID = 759832
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not inherently mean the customer is more or less likely to churn.&lt;/p&gt;

&lt;p&gt;The ID exists to identify the row, not to describe the customer.&lt;/p&gt;

&lt;p&gt;Important distinction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A column can be useful for identifying a record without being useful for predicting the target.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  26. But don't blindly drop columns
&lt;/h1&gt;

&lt;p&gt;This is an important improvement to your original reasoning.&lt;/p&gt;

&lt;p&gt;A column should not be dropped merely because it is an object/string.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract
InternetService
PaymentMethod
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;are categorical strings, but they can contain very useful predictive information.&lt;/p&gt;

&lt;p&gt;So the correct question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does this column contain useful information that is available at prediction time?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is this column numeric?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  27. Understanding your graphs
&lt;/h1&gt;

&lt;p&gt;You created many graphs. The important thing is to understand what question each graph answers.&lt;/p&gt;




&lt;h1&gt;
  
  
  28. Countplot
&lt;/h1&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A countplot answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many observations are in each category?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No churn → number of customers
Churn    → number of customers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This immediately helps you understand class balance.&lt;/p&gt;

&lt;p&gt;If the graph looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No Churn ███████████████████
Churn    ███████
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the classes are imbalanced.&lt;/p&gt;

&lt;p&gt;That matters because a model could achieve high accuracy simply by favoring the majority class.&lt;/p&gt;




&lt;h1&gt;
  
  
  29. Countplot with &lt;code&gt;hue&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Contract&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;hue&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How does the target category vary across another category?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract type
     ↓
Churn / No churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If month-to-month customers show much more churn than two-year customers, that is an important pattern worth investigating.&lt;/p&gt;

&lt;p&gt;But remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A graph showing association does not automatically prove causation.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  30. Histogram
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;figsize&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;histplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MonthlyCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;kde&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A histogram shows the distribution of a numerical variable.&lt;/p&gt;

&lt;p&gt;It helps answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where are most values?&lt;/li&gt;
&lt;li&gt;Is the data symmetric?&lt;/li&gt;
&lt;li&gt;Is it skewed?&lt;/li&gt;
&lt;li&gt;Are there multiple groups?&lt;/li&gt;
&lt;li&gt;Are there extreme values?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;frequency
   ^
   |       ███
   |     ███████
   |   █████████
   | ███████████
   +-----------------&amp;gt; value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  31. Why distributions matter
&lt;/h1&gt;

&lt;p&gt;Suppose a feature looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Most values: 0–100
A few values: 10,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That may indicate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;genuine outliers&lt;/li&gt;
&lt;li&gt;a different population&lt;/li&gt;
&lt;li&gt;data-entry errors&lt;/li&gt;
&lt;li&gt;a highly skewed distribution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A model may react differently depending on the algorithm.&lt;/p&gt;

&lt;p&gt;This is one reason EDA happens before modeling.&lt;/p&gt;




&lt;h1&gt;
  
  
  32. Boxplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A boxplot summarizes a distribution.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;       outlier
          •
          |
      ┌───────┐
      │       │
------│  box  │------
      │       │
      └───────┘
          |
       outlier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The box represents the middle portion of the data.&lt;/p&gt;

&lt;p&gt;The line inside the box is the median.&lt;/p&gt;

&lt;p&gt;Points outside the whiskers may be treated as potential outliers.&lt;/p&gt;

&lt;p&gt;Important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An outlier is not automatically an error.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A high-income customer may be perfectly legitimate.&lt;/p&gt;




&lt;h1&gt;
  
  
  33. Boxplot: target vs category
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Coverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a powerful graph.&lt;/p&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How does the distribution of Customer Lifetime Value differ between categories?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Coverage A → CLV distribution
Coverage B → CLV distribution
Coverage C → CLV distribution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;spread&lt;/li&gt;
&lt;li&gt;outliers&lt;/li&gt;
&lt;li&gt;overlap&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This can help you identify potentially useful relationships.&lt;/p&gt;




&lt;h1&gt;
  
  
  34. Scatterplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A scatterplot examines the relationship between two numerical variables.&lt;/p&gt;

&lt;p&gt;Each dot represents an observation.&lt;/p&gt;

&lt;p&gt;It helps answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"As X changes, does Y appear to change?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Possible patterns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Positive relationship:
  •
    •
      •
        •

Negative relationship:
        •
      •
    •
  •

No obvious relationship:
 •    •
    •
  •      •
     •
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pattern does not have to be a straight line.&lt;/p&gt;




&lt;h1&gt;
  
  
  35. Pairplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pairplot&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A pairplot shows many numerical relationships at once.&lt;/p&gt;

&lt;p&gt;It is useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;spotting correlations&lt;/li&gt;
&lt;li&gt;identifying clusters&lt;/li&gt;
&lt;li&gt;seeing distributions&lt;/li&gt;
&lt;li&gt;finding obvious relationships&lt;/li&gt;
&lt;li&gt;detecting possible outliers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The downside is that it becomes difficult to read when there are many columns.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Pairplot is excellent for small-to-medium exploratory datasets, but not something you blindly run on every large dataset.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  36. Correlation heatmap
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;heatmap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MonthlyCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]].&lt;/span&gt;&lt;span class="nf"&gt;corr&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;annot&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Correlation measures how two numerical variables move together in a linear relationship.&lt;/p&gt;

&lt;p&gt;A correlation close to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+1 → strong positive linear relationship
 0 → little/no linear relationship
-1 → strong negative linear relationship
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tenure ↑
TotalCharges ↑
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;could result in positive correlation.&lt;/p&gt;

&lt;p&gt;But correlation has an important limitation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correlation does not prove causation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Also, zero correlation does not necessarily mean "no relationship" because the relationship might be non-linear.&lt;/p&gt;




&lt;h1&gt;
  
  
  37. A key lesson from visualization
&lt;/h1&gt;

&lt;p&gt;The graphs answer different questions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Graph&lt;/th&gt;
&lt;th&gt;Main question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Countplot&lt;/td&gt;
&lt;td&gt;How many observations are in each category?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot + hue&lt;/td&gt;
&lt;td&gt;How does one category vary with another?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Histogram&lt;/td&gt;
&lt;td&gt;What does a numerical distribution look like?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boxplot&lt;/td&gt;
&lt;td&gt;What are the median, spread and potential outliers?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boxplot + category&lt;/td&gt;
&lt;td&gt;How does a numerical distribution differ by category?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scatterplot&lt;/td&gt;
&lt;td&gt;How do two numerical variables relate?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pairplot&lt;/td&gt;
&lt;td&gt;What relationships exist among several numerical variables?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heatmap&lt;/td&gt;
&lt;td&gt;How strongly are numerical variables linearly correlated?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The goal is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Create lots of graphs."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Use the right graph to answer the right question."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  38. Encoding: why does ML need it?
&lt;/h1&gt;

&lt;p&gt;Many ML algorithms work with numbers.&lt;/p&gt;

&lt;p&gt;But real-world data contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Male
Female

Yes
No

Month-to-month
One year
Two year

Fiber optic
DSL
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model needs these values represented numerically.&lt;/p&gt;

&lt;p&gt;This is called &lt;strong&gt;encoding&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  39. Label Encoding
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.preprocessing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LabelEncoder&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and encoded categorical columns.&lt;/p&gt;

&lt;p&gt;For a binary column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can be reasonable for a binary variable.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn:
No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes intuitive sense.&lt;/p&gt;




&lt;h1&gt;
  
  
  40. Why LabelEncoder can be dangerous for multi-category variables
&lt;/h1&gt;

&lt;p&gt;Suppose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract:
Month-to-month
One year
Two year
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Label encoding might produce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Month-to-month → 0
One year       → 1
Two year       → 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A numerical model may interpret that as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Two year &amp;gt; One year &amp;gt; Month-to-month
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That ordering may not be appropriate for the algorithm.&lt;/p&gt;

&lt;p&gt;The numbers are just codes.&lt;/p&gt;

&lt;p&gt;They do not automatically mean that category 2 is "twice" category 1.&lt;/p&gt;

&lt;p&gt;For nominal categories, one-hot encoding is usually safer.&lt;/p&gt;




&lt;h1&gt;
  
  
  41. One-hot encoding
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_dummies&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;drop_first&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This transforms categories into separate binary columns.&lt;/p&gt;

&lt;p&gt;For:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;InternetService:
DSL
Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we might create columns such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;InternetService_Fiber optic
InternetService_No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with the remaining category represented by both values being 0.&lt;/p&gt;

&lt;p&gt;This avoids inventing a fake numerical ordering.&lt;/p&gt;




&lt;h1&gt;
  
  
  42. Why &lt;code&gt;drop_first=True&lt;/code&gt;?
&lt;/h1&gt;

&lt;p&gt;If a categorical variable has several categories, using all one-hot columns can introduce redundancy in some models.&lt;/p&gt;

&lt;p&gt;Dropping one category creates a reference category.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DSL
Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;could become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fiber optic = 0
No = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the reference category:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DSL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For beginner understanding, the key idea is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One-hot encoding converts categories into machine-readable indicator variables without pretending the categories have numerical order.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  43. Train/test split
&lt;/h1&gt;

&lt;p&gt;You repeatedly used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the most important ML concepts.&lt;/p&gt;

&lt;p&gt;Suppose we have 1,000 examples.&lt;/p&gt;

&lt;p&gt;We might use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;700 → training
300 → testing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The training set is used to learn the model.&lt;/p&gt;

&lt;p&gt;The test set is held back to evaluate how the trained model performs on unseen data.&lt;/p&gt;

&lt;p&gt;Think of it like an exam:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data = practice questions
Test data     = unseen exam questions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  44. Why not train and test on the same data?
&lt;/h1&gt;

&lt;p&gt;Because the model could simply memorize the training examples.&lt;/p&gt;

&lt;p&gt;Suppose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 99%
Test score     = 65%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a warning sign.&lt;/p&gt;

&lt;p&gt;The model learned the training data extremely well but does not generalize well.&lt;/p&gt;

&lt;p&gt;This is called:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Overfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  45. Underfitting
&lt;/h1&gt;

&lt;p&gt;The opposite can happen.&lt;/p&gt;

&lt;p&gt;If the model is too simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 60%
Test score     = 58%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It may not have learned enough from the data.&lt;/p&gt;

&lt;p&gt;This is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Underfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A useful mental picture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Underfitting
Model too simple
       ↓
misses important patterns

Good fit
Learns useful patterns
       ↓
works on unseen data

Overfitting
Model learns noise/details
       ↓
great training performance
poor unseen performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  46. Why &lt;code&gt;random_state&lt;/code&gt; matters
&lt;/h1&gt;

&lt;p&gt;In one Telco notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the improved notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A random state makes the split reproducible.&lt;/p&gt;

&lt;p&gt;Without it, the random split can change between runs.&lt;/p&gt;

&lt;p&gt;That means your score may change.&lt;/p&gt;

&lt;p&gt;With:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you can reproduce the same split.&lt;/p&gt;

&lt;p&gt;The number &lt;code&gt;42&lt;/code&gt; is not magical.&lt;/p&gt;

&lt;p&gt;Any fixed integer can serve this purpose.&lt;/p&gt;




&lt;h1&gt;
  
  
  47. Why &lt;code&gt;stratify=Y&lt;/code&gt; is useful for classification
&lt;/h1&gt;

&lt;p&gt;In your improved Telco notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;stratify&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This helps preserve the class distribution between training and testing data.&lt;/p&gt;

&lt;p&gt;For example, if the full dataset contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;73% No Churn
27% Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;stratification aims to keep roughly the same proportion in both splits.&lt;/p&gt;

&lt;p&gt;This is particularly useful when classes are imbalanced.&lt;/p&gt;




&lt;h1&gt;
  
  
  48. Logistic Regression
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.linear_model&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LogisticRegression&lt;/span&gt;

&lt;span class="n"&gt;model_lr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LogisticRegression&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;model_lr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Despite its name, Logistic Regression is commonly used for &lt;strong&gt;classification&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Its goal is to estimate the probability of belonging to a class.&lt;/p&gt;

&lt;p&gt;For binary classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;features
   ↓
Logistic Regression
   ↓
probability
   ↓
class
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Probability of churn = 0.82
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model may classify that customer as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;depending on its decision threshold.&lt;/p&gt;




&lt;h1&gt;
  
  
  49. Why Logistic Regression is a good beginner model
&lt;/h1&gt;

&lt;p&gt;It is useful because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;relatively simple&lt;/li&gt;
&lt;li&gt;fast&lt;/li&gt;
&lt;li&gt;often strong as a baseline&lt;/li&gt;
&lt;li&gt;easier to interpret than many complex models&lt;/li&gt;
&lt;li&gt;naturally suited to binary classification&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A good ML habit is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Start with a simple baseline before reaching for complex models.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  50. Support Vector Machine / SVC
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.svm&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SVC&lt;/span&gt;

&lt;span class="n"&gt;model_lr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SVC&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The variable name &lt;code&gt;model_lr&lt;/code&gt; is misleading here.&lt;/p&gt;

&lt;p&gt;It is actually an SVC model.&lt;/p&gt;

&lt;p&gt;SVC means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Support Vector Classifier&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The basic idea is to find a decision boundary that separates classes.&lt;/p&gt;

&lt;p&gt;In simple 2D data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Class A: ● ● ●

--------- decision boundary ---------

Class B: ▲ ▲ ▲
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SVM tries to find a boundary with a useful margin between classes.&lt;/p&gt;




&lt;h1&gt;
  
  
  51. Why scaling matters for SVM
&lt;/h1&gt;

&lt;p&gt;SVM is sensitive to feature scale.&lt;/p&gt;

&lt;p&gt;Suppose one feature ranges from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0–1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and another ranges from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0–100,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The larger-scale feature can dominate distance-related calculations.&lt;/p&gt;

&lt;p&gt;That is why scaling is often important for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SVM&lt;/li&gt;
&lt;li&gt;KNN&lt;/li&gt;
&lt;li&gt;Logistic Regression in many situations&lt;/li&gt;
&lt;li&gt;neural networks&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  52. Decision Tree
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeClassifier&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A decision tree makes predictions using a sequence of questions.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is tenure &amp;lt; 12?
       /       \
     Yes       No
     /           \
Is charge &amp;gt; X?   ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eventually the tree reaches a leaf containing a prediction.&lt;/p&gt;

&lt;p&gt;This is attractive because it is easy to visualize conceptually.&lt;/p&gt;




&lt;h1&gt;
  
  
  53. Why trees can overfit
&lt;/h1&gt;

&lt;p&gt;A tree can keep splitting the data until it creates extremely specific rules.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If feature A &amp;gt; 10
and feature B &amp;lt; 4
and feature C = 7
and feature D &amp;gt; 92
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eventually the tree may memorize training examples.&lt;/p&gt;

&lt;p&gt;That is why you later tried:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;
&lt;span class="n"&gt;min_samples_split&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;
&lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are ways of controlling tree complexity.&lt;/p&gt;




&lt;h1&gt;
  
  
  54. Random Forest
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;RandomForestRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A random forest combines many decision trees.&lt;/p&gt;

&lt;p&gt;Instead of relying on one tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tree 1 ─┐
Tree 2 ─┤
Tree 3 ─┤
Tree 4 ─┤ → combined prediction
...     ┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The idea is that many different trees can produce a more robust prediction than a single tree.&lt;/p&gt;

&lt;p&gt;This is an example of an &lt;strong&gt;ensemble method&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  55. Bagging
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;BaggingRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_samples&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Bagging means combining models trained on different samples of the data.&lt;/p&gt;

&lt;p&gt;The overall idea:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dataset
  ↓
Different samples
  ↓
Multiple models
  ↓
Combine predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is often to reduce variance and make predictions more stable.&lt;/p&gt;

&lt;p&gt;Random Forest can be viewed as a specialized tree-based ensemble that adds additional randomness in how trees are built.&lt;/p&gt;




&lt;h1&gt;
  
  
  56. AdaBoost
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;AdaBoostRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AdaBoost means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Adaptive Boosting&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of training many independent models and simply averaging them, boosting builds models sequentially.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model 1
  ↓
find mistakes
  ↓
Model 2 focuses more on difficult cases
  ↓
Model 3 focuses further
  ↓
combine models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key idea:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Later models try to improve on the weaknesses of earlier models.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  57. K-Nearest Neighbors (KNN)
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;KNeighborsRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;KNN predicts using nearby examples.&lt;/p&gt;

&lt;p&gt;Imagine a new customer.&lt;/p&gt;

&lt;p&gt;The algorithm asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which existing customers are most similar to this customer?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then it uses those neighbors to estimate the output.&lt;/p&gt;

&lt;p&gt;For regression, the prediction is commonly based on the neighbors' target values.&lt;/p&gt;




&lt;h1&gt;
  
  
  58. Why KNN needs scaling
&lt;/h1&gt;

&lt;p&gt;KNN relies on distance.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Income:   0–100,000
Age:      18–80
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If raw values are used, income can dominate the distance calculation.&lt;/p&gt;

&lt;p&gt;Scaling puts features onto comparable scales.&lt;/p&gt;

&lt;p&gt;That is why you experimented with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  59. StandardScaler
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;X_train_scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;X_test_scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an important pattern.&lt;/p&gt;

&lt;p&gt;Notice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;The scaler learns the training data's statistics.&lt;/p&gt;

&lt;p&gt;Then the exact same transformation is applied to the test data.&lt;/p&gt;

&lt;p&gt;You should not independently fit the scaler on the test set.&lt;/p&gt;

&lt;p&gt;Otherwise information from the test set leaks into preprocessing.&lt;/p&gt;

&lt;p&gt;This idea is called avoiding:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Data leakage&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  60. Data leakage
&lt;/h1&gt;

&lt;p&gt;Data leakage happens when information that should not be available to the model during training accidentally influences the model.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Train preprocessing
     ↓
uses information from test data
     ↓
test is no longer truly unseen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That can make evaluation look better than real-world performance.&lt;/p&gt;

&lt;p&gt;The correct principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fit preprocessing steps using training data, then apply them to validation/test/new data.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A production ML pipeline should ideally bundle preprocessing and modeling together.&lt;/p&gt;




&lt;h1&gt;
  
  
  61. Linear Regression
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;LinearRegression&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for Customer Lifetime Value.&lt;/p&gt;

&lt;p&gt;Linear regression tries to model a relationship like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = b0 + b1X1 + b2X2 + ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For one feature:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = intercept + slope × X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is to find coefficients that produce predictions close to the observed values.&lt;/p&gt;




&lt;h1&gt;
  
  
  62. Regression score: what does &lt;code&gt;.score()&lt;/code&gt; mean?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the scikit-learn regression models you used, &lt;code&gt;.score()&lt;/code&gt; generally returns &lt;strong&gt;R²&lt;/strong&gt;, the coefficient of determination.&lt;/p&gt;

&lt;p&gt;Very roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R² = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the model explains the observed variation perfectly on that dataset.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R² = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the model is no better than a simple baseline based on predicting the mean target.&lt;/p&gt;

&lt;p&gt;R² can also be negative on unseen data.&lt;/p&gt;

&lt;p&gt;Important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;R² is not the same thing as accuracy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For regression, do not say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"My regression model has 90% accuracy"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;just because R² is 0.90.&lt;/p&gt;

&lt;p&gt;Instead say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The model achieved an R² of 0.90."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  63. Classification accuracy
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;accuracy_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test_pred&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Accuracy is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;correct predictions
-------------------
total predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;90 correct
100 total
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;gives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;90% accuracy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Accuracy is easy to understand.&lt;/p&gt;

&lt;p&gt;But it can be misleading when classes are highly imbalanced.&lt;/p&gt;




&lt;h1&gt;
  
  
  64. Confusion matrix
&lt;/h1&gt;

&lt;p&gt;Your later notes/workflow refer to concepts such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TP
TN
FP
FN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A confusion matrix organizes classification predictions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Actual
                 No       Yes
Predicted No     TN       FN
Predicted Yes    FP       TP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meaning:&lt;/p&gt;

&lt;h3&gt;
  
  
  True Positive
&lt;/h3&gt;

&lt;p&gt;Model predicted positive and it was positive.&lt;/p&gt;

&lt;h3&gt;
  
  
  True Negative
&lt;/h3&gt;

&lt;p&gt;Model predicted negative and it was negative.&lt;/p&gt;

&lt;h3&gt;
  
  
  False Positive
&lt;/h3&gt;

&lt;p&gt;Model predicted positive but it was negative.&lt;/p&gt;

&lt;h3&gt;
  
  
  False Negative
&lt;/h3&gt;

&lt;p&gt;Model predicted negative but it was positive.&lt;/p&gt;

&lt;p&gt;This is much more informative than accuracy alone when the cost of errors differs.&lt;/p&gt;




&lt;h1&gt;
  
  
  65. Precision and Recall
&lt;/h1&gt;

&lt;p&gt;For classification:&lt;/p&gt;

&lt;h3&gt;
  
  
  Precision
&lt;/h3&gt;

&lt;p&gt;Of everything the model predicted as positive:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many were actually positive?&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Precision = TP / (TP + FP)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Recall
&lt;/h3&gt;

&lt;p&gt;Of everything that was actually positive:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many did the model find?&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recall = TP / (TP + FN)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right metric depends on the problem.&lt;/p&gt;

&lt;p&gt;For churn prediction, for example, missing a customer who is about to churn may be more important than contacting an extra customer who would have stayed.&lt;/p&gt;




&lt;h1&gt;
  
  
  66. Your Telco Customer Churn project
&lt;/h1&gt;

&lt;p&gt;This is probably the clearest classification example in your notebooks.&lt;/p&gt;

&lt;p&gt;The business question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can we predict whether a telecom customer will churn?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The target is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Features are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer data
     ↓
Clean data
     ↓
Explore data
     ↓
Encode categories
     ↓
Split train/test
     ↓
Train classifier
     ↓
Predict churn
     ↓
Evaluate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  67. Telco: removing customerID
&lt;/h1&gt;

&lt;p&gt;You did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customerID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is reasonable because the ID is an identifier rather than a customer characteristic.&lt;/p&gt;

&lt;p&gt;You also commented that it could be kept separately if you later wanted to map predictions back to customers.&lt;/p&gt;

&lt;p&gt;That is a good practical idea.&lt;/p&gt;

&lt;p&gt;In production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID
   ↓
keep separately for business reporting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID
   X
do not necessarily use as a model feature
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  68. Telco: fixing TotalCharges
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coerce&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;errors="coerce"&lt;/code&gt; means invalid values are converted to &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Then you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;fillna&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in one notebook and median imputation in another.&lt;/p&gt;

&lt;p&gt;These are two different decisions.&lt;/p&gt;

&lt;p&gt;The important lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Missing-value handling should be based on the meaning of the data, not simply on whichever replacement is easiest.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example, if missing &lt;code&gt;TotalCharges&lt;/code&gt; occurs because a customer has zero tenure and has just joined, then filling with zero may have a meaningful business interpretation.&lt;/p&gt;

&lt;p&gt;If missingness represents an unknown measurement, median imputation may be more appropriate.&lt;/p&gt;




&lt;h1&gt;
  
  
  69. Telco: the first modeling version
&lt;/h1&gt;

&lt;p&gt;The earlier Telco notebook used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;LabelEncoder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for many categorical columns.&lt;/p&gt;

&lt;p&gt;This is okay as a learning exercise for understanding encoding, but it is not ideal for every categorical variable.&lt;/p&gt;

&lt;p&gt;A better version in your later notebook used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_dummies&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;multi_cols&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;drop_first&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a meaningful improvement.&lt;/p&gt;




&lt;h1&gt;
  
  
  70. Telco: your improved preprocessing pipeline
&lt;/h1&gt;

&lt;p&gt;The later notebook did something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw data
   ↓
Drop customerID
   ↓
Convert TotalCharges
   ↓
Encode binary columns
   ↓
One-hot encode multi-category columns
   ↓
Separate X and Y
   ↓
Scale X
   ↓
Train Logistic Regression
   ↓
Create Gradio interface
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is much closer to a real ML application.&lt;/p&gt;




&lt;h1&gt;
  
  
  71. Gradio: turning a model into an application
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;gradio&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;gr&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and built a UI.&lt;/p&gt;

&lt;p&gt;This is an important step because it changes the project from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Notebook experiment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User input
    ↓
Preprocessing
    ↓
Model
    ↓
Prediction
    ↓
Human-readable result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gender: Female
Tenure: 12 months
Contract: Month-to-month
Monthly charges: $70
...
          ↓
    Predict Churn
          ↓
High Churn Risk
Probability: ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This demonstrates an important ML engineering concept:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A model is useful only when it can be integrated into a workflow where people or systems can use its predictions.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  72. Very important: preprocessing must match training
&lt;/h1&gt;

&lt;p&gt;Your Gradio code contains a good lesson.&lt;/p&gt;

&lt;p&gt;During training you:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;encode the categorical values&lt;/li&gt;
&lt;li&gt;create dummy columns&lt;/li&gt;
&lt;li&gt;scale the features&lt;/li&gt;
&lt;li&gt;train the model&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;During prediction you must do the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;raw user input
    ↓
same encoding
    ↓
same columns
    ↓
same scaling
    ↓
model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the training data has:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature_1
feature_2
feature_3
feature_4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but the application sends:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature_1
feature_3
feature_4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model cannot interpret the input correctly.&lt;/p&gt;

&lt;p&gt;Your use of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;input_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reindex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;feature_columns&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;fill_value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is intended to make the input columns match the training feature structure.&lt;/p&gt;

&lt;p&gt;That is an important practical idea.&lt;/p&gt;




&lt;h1&gt;
  
  
  73. A stronger production pattern: Pipeline
&lt;/h1&gt;

&lt;p&gt;A cleaner production approach is often:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;Pipeline&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preprocessor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;LogisticRegression&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then preprocessing and prediction are tied together.&lt;/p&gt;

&lt;p&gt;This reduces the risk of accidentally applying different transformations during training and inference.&lt;/p&gt;

&lt;p&gt;Your notebook demonstrates the concept manually, which is useful for learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  74. Your Auto Insurance / Customer Lifetime Value project
&lt;/h1&gt;

&lt;p&gt;The &lt;code&gt;13-Aug.ipynb&lt;/code&gt; notebook moves into regression.&lt;/p&gt;

&lt;p&gt;The target is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This changes the question from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Will this customer churn?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What numerical Customer Lifetime Value should we predict?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Therefore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Classification → Churn
Regression     → Customer Lifetime Value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  75. Insurance dataset: initial exploration
&lt;/h1&gt;

&lt;p&gt;You loaded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AutoInsurance.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then inspected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You also removed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer
Effective To Date
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from the modeling DataFrame.&lt;/p&gt;

&lt;p&gt;Again, the purpose is to separate identifiers / dates that were not being used in the current modeling approach.&lt;/p&gt;

&lt;p&gt;However, an important ML lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Dates are not automatically useless.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A date may contain useful information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;month&lt;/li&gt;
&lt;li&gt;day of week&lt;/li&gt;
&lt;li&gt;season&lt;/li&gt;
&lt;li&gt;year&lt;/li&gt;
&lt;li&gt;time since an event&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A stronger project could engineer useful date features instead of simply dropping every date column.&lt;/p&gt;




&lt;h1&gt;
  
  
  76. Insurance EDA
&lt;/h1&gt;

&lt;p&gt;You explored:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;State&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Coverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;hist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a good progression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Categorical distribution
        ↓
Numerical distribution
        ↓
Numerical vs categorical
        ↓
Numerical vs numerical
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You are gradually asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What might explain the target?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  77. Automatically finding numerical and categorical columns
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;numerical_cols&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_dtypes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;float64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;

&lt;span class="n"&gt;categorical_cols&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_dtypes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful because instead of manually listing every column, Python can identify columns based on data type.&lt;/p&gt;

&lt;p&gt;Then you looped through them to create graphs.&lt;/p&gt;

&lt;p&gt;This is the beginning of writing reusable data-analysis code.&lt;/p&gt;




&lt;h1&gt;
  
  
  78. Automated boxplots
&lt;/h1&gt;

&lt;p&gt;You created boxplots for all numerical columns.&lt;/p&gt;

&lt;p&gt;This is useful because you can quickly inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;distributions&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;spread&lt;/li&gt;
&lt;li&gt;potential outliers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, remember that a large collection of graphs can become overwhelming.&lt;/p&gt;

&lt;p&gt;The goal should eventually move from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Let's plot everything."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Let's plot the variables that help answer our question."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a sign of growing from beginner EDA toward professional analysis.&lt;/p&gt;




&lt;h1&gt;
  
  
  79. Binning numerical variables for boxplots
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cut&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;bins&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This divides a numerical variable into five intervals.&lt;/p&gt;

&lt;p&gt;Then you compare CLV across those ranges.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Income range 1 → CLV distribution
Income range 2 → CLV distribution
Income range 3 → CLV distribution
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful for visualization because a boxplot expects categories on one axis.&lt;/p&gt;

&lt;p&gt;But the number and boundaries of bins can affect the story you see.&lt;/p&gt;

&lt;p&gt;So binning is mainly a visualization technique here, not necessarily something you should automatically use as a model feature.&lt;/p&gt;




&lt;h1&gt;
  
  
  80. Comparing many regression algorithms
&lt;/h1&gt;

&lt;p&gt;One of the biggest learning steps in your final notebook was trying multiple regressors:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Linear Regression
KNN Regressor
SVR
Decision Tree Regressor
Bagging Regressor
AdaBoost Regressor
Random Forest Regressor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is valuable because it teaches a core ML lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Different algorithms make different assumptions and capture different types of patterns.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is no universally best algorithm.&lt;/p&gt;




&lt;h1&gt;
  
  
  81. Linear Regression vs tree-based models
&lt;/h1&gt;

&lt;p&gt;Linear Regression assumes a relationship that can be represented through a linear combination of features.&lt;/p&gt;

&lt;p&gt;Tree-based methods can represent more complex non-linear relationships.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Linear model:

Y
│       /
│     /
│   /
│ /
└──────── X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A tree-based model can approximate more irregular patterns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y
│    ┌────
│    │
│ ───┘
│
└──────── X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why trying multiple model families can be useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  82. Why model comparison matters
&lt;/h1&gt;

&lt;p&gt;Suppose you obtain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model             Train R²    Test R²
Linear Regression   0.70       0.68
KNN                 0.90       0.62
Decision Tree       0.99       0.55
Random Forest       0.91       0.78
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should not automatically choose the model with the highest training score.&lt;/p&gt;

&lt;p&gt;The more important question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which model generalizes well to unseen data?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here Random Forest would be more interesting than the tree with 0.99 training R².&lt;/p&gt;




&lt;h1&gt;
  
  
  83. Log transformation of Customer Lifetime Value
&lt;/h1&gt;

&lt;p&gt;You later used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because some target variables are heavily skewed.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Most CLV values: relatively small
A few CLV values: extremely large
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model may have difficulty fitting such a distribution.&lt;/p&gt;

&lt;p&gt;A log transformation compresses large values.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original:
1
10
100
1000
10000

Log scale:
small differences between large values
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can make the target distribution easier for some models to learn.&lt;/p&gt;




&lt;h1&gt;
  
  
  84. Reversing the log transformation
&lt;/h1&gt;

&lt;p&gt;After predicting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_actual&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expm1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;If:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y_log = log(1 + Y)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = exp(Y_log) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;np.expm1()&lt;/code&gt; performs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exp(x) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;log target
   ↓
model
   ↓
log prediction
   ↓
expm1
   ↓
original target scale
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  85. An important issue in the notebook: split consistency
&lt;/h1&gt;

&lt;p&gt;In your log-transform section, you did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;X_train_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_test_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Y_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is okay as a new split, but it means your earlier train/test split and later train/test split are not necessarily the same.&lt;/p&gt;

&lt;p&gt;For a clean experiment, define the split once and reuse it consistently.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;y_train_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;y_test_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes comparisons easier to reason about.&lt;/p&gt;




&lt;h1&gt;
  
  
  86. Another important issue: predicting on the full dataset
&lt;/h1&gt;

&lt;p&gt;You later did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model_rf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and saved predictions for the full dataset.&lt;/p&gt;

&lt;p&gt;That can be useful for producing a business-facing prediction file.&lt;/p&gt;

&lt;p&gt;But these are &lt;strong&gt;not unbiased test predictions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because &lt;code&gt;X&lt;/code&gt; includes training observations the model already saw.&lt;/p&gt;

&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Predictions on X_test
→ useful for evaluating unseen data

Predictions on X
→ useful for generating predictions for all records,
   but not for measuring generalization
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction is very important.&lt;/p&gt;




&lt;h1&gt;
  
  
  87. Hyperparameters
&lt;/h1&gt;

&lt;p&gt;Later you changed models from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_split&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These settings are called &lt;strong&gt;hyperparameters&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;They are chosen by us before training.&lt;/p&gt;

&lt;p&gt;The model learns its internal parameters from the training data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parameters
&lt;/h3&gt;

&lt;p&gt;Learned by the algorithm.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hyperparameters
&lt;/h3&gt;

&lt;p&gt;Set by us.&lt;/p&gt;

&lt;p&gt;This distinction is fundamental.&lt;/p&gt;




&lt;h1&gt;
  
  
  88. Understanding your tree hyperparameters
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;max_depth&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Controls how deep the tree can grow.&lt;/p&gt;

&lt;p&gt;Smaller:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;simpler tree
less overfitting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Larger:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;more complex tree
greater overfitting risk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  &lt;code&gt;min_samples_split&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Minimum number of samples required before a node can be split.&lt;/p&gt;

&lt;p&gt;Higher values make the tree more conservative.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;min_samples_leaf&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Minimum number of samples allowed in a leaf.&lt;/p&gt;

&lt;p&gt;Higher values prevent extremely tiny leaves.&lt;/p&gt;




&lt;h1&gt;
  
  
  89. Random Forest hyperparameters
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;RandomForestRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;code&gt;n_estimators=200&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Build 200 trees.&lt;/p&gt;

&lt;p&gt;More trees can make the ensemble more stable, though computation increases.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;max_depth=8&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Limits tree complexity.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;min_samples_leaf=5&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Prevents leaves from becoming too specific.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;random_state=42&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Makes the experiment reproducible.&lt;/p&gt;




&lt;h1&gt;
  
  
  90. Bagging hyperparameters
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;BaggingRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_samples&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This means, conceptually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;build many estimators&lt;/li&gt;
&lt;li&gt;each estimator sees a sample of the training rows&lt;/li&gt;
&lt;li&gt;each estimator can use a subset of features&lt;/li&gt;
&lt;li&gt;combine their predictions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Again, the goal is to create a robust ensemble rather than rely on one model.&lt;/p&gt;




&lt;h1&gt;
  
  
  91. Scaling experiment in your notebook
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and compared KNN/SVR.&lt;/p&gt;

&lt;p&gt;This is a very useful experiment because it demonstrates:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Preprocessing requirements depend on the algorithm.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tree-based algorithms generally do not need feature scaling in the same way that distance- or margin-based methods do.&lt;/p&gt;

&lt;p&gt;KNN and SVM are much more sensitive to feature scale.&lt;/p&gt;




&lt;h1&gt;
  
  
  92. A subtle issue in your scaling experiment
&lt;/h1&gt;

&lt;p&gt;You created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train_scaled&lt;/span&gt;
&lt;span class="n"&gt;X_test_scaled&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but your KNN code first trained on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train_log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before later looping over scaled KNN.&lt;/p&gt;

&lt;p&gt;That is useful as experimentation, but for a clean article you should present the comparison more systematically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KNN without scaling
        vs
KNN with scaling
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and explicitly explain why the result changed.&lt;/p&gt;

&lt;p&gt;For SVR, you correctly used the scaled data in the shown section.&lt;/p&gt;




&lt;h1&gt;
  
  
  93. The complete mental model
&lt;/h1&gt;

&lt;p&gt;After these five days, the entire workflow can be remembered as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             REAL-WORLD QUESTION
                     ↓
             Define the target
                     ↓
              Collect data
                     ↓
              Understand data
                     ↓
          ┌──────────┴──────────┐
          ↓                     ↓
      Numerical             Categorical
          ↓                     ↓
      distributions        category counts
          └──────────┬──────────┘
                     ↓
                    EDA
                     ↓
              Clean the data
                     ↓
          Handle missing values
                     ↓
             Encode categories
                     ↓
          Select useful features
                     ↓
              Split train/test
                     ↓
          Scale when appropriate
                     ↓
              Train baseline
                     ↓
          Evaluate on test data
                     ↓
             Try other models
                     ↓
         Tune hyperparameters
                     ↓
       Select based on validation
                     ↓
             Final evaluation
                     ↓
              Make predictions
                     ↓
       Integrate into an application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  94. What each notebook taught you
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Day 1 — Auto MPG
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;load a dataset&lt;/li&gt;
&lt;li&gt;inspect rows and columns&lt;/li&gt;
&lt;li&gt;understand shape and data types&lt;/li&gt;
&lt;li&gt;inspect summary statistics&lt;/li&gt;
&lt;li&gt;identify categorical/text columns&lt;/li&gt;
&lt;li&gt;detect messy numerical values&lt;/li&gt;
&lt;li&gt;visualize distributions&lt;/li&gt;
&lt;li&gt;use countplots&lt;/li&gt;
&lt;li&gt;use histograms&lt;/li&gt;
&lt;li&gt;use pairplots&lt;/li&gt;
&lt;li&gt;understand missing values&lt;/li&gt;
&lt;li&gt;convert &lt;code&gt;"?"&lt;/code&gt; to &lt;code&gt;NaN&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;convert text to numerical data&lt;/li&gt;
&lt;li&gt;use median imputation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Before modeling, understand and clean your data.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 2 — Telco Customer Churn
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identify a classification target&lt;/li&gt;
&lt;li&gt;remove an identifier&lt;/li&gt;
&lt;li&gt;handle &lt;code&gt;TotalCharges&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;visualize class balance&lt;/li&gt;
&lt;li&gt;compare categories against churn&lt;/li&gt;
&lt;li&gt;use boxplots and histograms&lt;/li&gt;
&lt;li&gt;use pairplots&lt;/li&gt;
&lt;li&gt;inspect correlations&lt;/li&gt;
&lt;li&gt;encode categorical data&lt;/li&gt;
&lt;li&gt;split into training/testing data&lt;/li&gt;
&lt;li&gt;train Logistic Regression&lt;/li&gt;
&lt;li&gt;train SVC&lt;/li&gt;
&lt;li&gt;train Decision Tree&lt;/li&gt;
&lt;li&gt;compare train/test performance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A classification model predicts categories, but the quality of the prediction depends heavily on preprocessing and evaluation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 3 — Improving the Telco project
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one-hot encode multi-category variables&lt;/li&gt;
&lt;li&gt;use &lt;code&gt;random_state&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;use &lt;code&gt;stratify&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;scale features&lt;/li&gt;
&lt;li&gt;train Logistic Regression&lt;/li&gt;
&lt;li&gt;build a prediction interface with Gradio&lt;/li&gt;
&lt;li&gt;reproduce preprocessing at inference time&lt;/li&gt;
&lt;li&gt;output a human-readable prediction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;An ML model is only one part of an ML application.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 4/5 — Customer Lifetime Value
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;formulate a regression problem&lt;/li&gt;
&lt;li&gt;inspect numerical/categorical relationships&lt;/li&gt;
&lt;li&gt;automate EDA&lt;/li&gt;
&lt;li&gt;encode categories&lt;/li&gt;
&lt;li&gt;split train/test&lt;/li&gt;
&lt;li&gt;compare multiple regressors&lt;/li&gt;
&lt;li&gt;understand R²&lt;/li&gt;
&lt;li&gt;try target transformation&lt;/li&gt;
&lt;li&gt;use log transformation&lt;/li&gt;
&lt;li&gt;reverse the transformation&lt;/li&gt;
&lt;li&gt;experiment with scaling&lt;/li&gt;
&lt;li&gt;tune hyperparameters&lt;/li&gt;
&lt;li&gt;use ensemble models&lt;/li&gt;
&lt;li&gt;generate prediction files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Different models capture different patterns, and model selection should be based on performance on unseen data—not training performance alone.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  95. What I would change before publishing the project as "production-quality"
&lt;/h1&gt;

&lt;p&gt;Your notebooks are excellent learning material because they show experimentation. For a public article, however, it is worth separating:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Learning experiments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recommended ML workflow
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here are the most important improvements.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 1 — Use clear train/test terminology
&lt;/h2&gt;

&lt;p&gt;Prefer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;rather than mixing &lt;code&gt;Y&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Python convention generally uses lowercase &lt;code&gt;y&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 2 — Make experiments reproducible
&lt;/h2&gt;

&lt;p&gt;Use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;consistently when appropriate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 3 — Use stratification for classification
&lt;/h2&gt;

&lt;p&gt;For churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stratify&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Improvement 4 — Avoid LabelEncoder for nominal multi-class input features
&lt;/h2&gt;

&lt;p&gt;Use one-hot encoding for categories such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PaymentMethod
Contract
InternetService
VehicleClass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Label encoding is much more naturally suited to targets or truly ordinal categories.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 5 — Fit preprocessing only on training data
&lt;/h2&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data
    ↓
fit scaler
    ↓
transform training

Test data
    ↓
transform using existing scaler
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not independently fit preprocessing on the test data.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 6 — Use proper evaluation metrics
&lt;/h2&gt;

&lt;p&gt;For classification, consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accuracy
Precision
Recall
F1
Confusion Matrix
ROC-AUC
PR-AUC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For regression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R²
MAE
MSE
RMSE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which metrics matter depends on the business problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 7 — Don't compare only training scores
&lt;/h2&gt;

&lt;p&gt;Always look at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training performance
+
Validation/test performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A very high training score with much lower test performance can indicate overfitting.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 8 — Keep a validation strategy for model selection
&lt;/h2&gt;

&lt;p&gt;A single test set should ideally be kept for final evaluation.&lt;/p&gt;

&lt;p&gt;During experimentation, use cross-validation or a validation set to choose models/hyperparameters.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data
   ↓
Cross-validation / validation
   ↓
Choose model
   ↓
Final test set
   ↓
Final unbiased-ish estimate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  96. A very important distinction: EDA vs ML
&lt;/h1&gt;

&lt;p&gt;EDA asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What does my data look like?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Machine learning asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can I learn a useful relationship that lets me make predictions on new data?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These overlap, but they are not the same.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not build a model.&lt;/p&gt;

&lt;p&gt;It helps you understand the data before deciding how to model it.&lt;/p&gt;




&lt;h1&gt;
  
  
  97. Another important distinction: correlation vs prediction
&lt;/h1&gt;

&lt;p&gt;A strong correlation can be useful.&lt;/p&gt;

&lt;p&gt;But:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;high correlation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not automatically mean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;excellent predictive model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prediction depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;multiple features&lt;/li&gt;
&lt;li&gt;data quality&lt;/li&gt;
&lt;li&gt;model type&lt;/li&gt;
&lt;li&gt;noise&lt;/li&gt;
&lt;li&gt;generalization&lt;/li&gt;
&lt;li&gt;preprocessing&lt;/li&gt;
&lt;li&gt;evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So don't judge an ML problem from one graph.&lt;/p&gt;




&lt;h1&gt;
  
  
  98. Another important distinction: model vs algorithm
&lt;/h1&gt;

&lt;p&gt;It is common to hear:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I trained a model."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The algorithm is the learning procedure.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Random Forest = algorithm/model family
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After training:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model_rf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you have a fitted model containing learned information.&lt;/p&gt;

&lt;p&gt;A useful mental model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Algorithm
   +
Training data
   ↓
Fitted model
   ↓
Predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  99. What "learning" actually means
&lt;/h1&gt;

&lt;p&gt;This is probably the most important concept to understand.&lt;/p&gt;

&lt;p&gt;When you run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model is not magically understanding the world.&lt;/p&gt;

&lt;p&gt;It is optimizing internal parameters so that its predictions match the training examples according to the algorithm's objective.&lt;/p&gt;

&lt;p&gt;For example, linear regression learns coefficients.&lt;/p&gt;

&lt;p&gt;A tree learns split rules.&lt;/p&gt;

&lt;p&gt;A neural network learns weights.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Training means finding model parameters that make the model perform well according to a defined objective.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  100. What happens during prediction?
&lt;/h1&gt;

&lt;p&gt;Once training is complete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prediction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model does not learn again from the test examples.&lt;/p&gt;

&lt;p&gt;It applies what it learned during training.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training:

X_train + y_train
        ↓
      model.fit()
        ↓
   learned model

Prediction:

X_test
   ↓
learned model
   ↓
prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  101. The biggest beginner misconception to avoid
&lt;/h1&gt;

&lt;p&gt;Do not think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I used Random Forest, so I did machine learning."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The real ML skill is being able to explain:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What problem are you solving?&lt;/li&gt;
&lt;li&gt;What is the target?&lt;/li&gt;
&lt;li&gt;What features are available?&lt;/li&gt;
&lt;li&gt;What does the data look like?&lt;/li&gt;
&lt;li&gt;What problems did you find?&lt;/li&gt;
&lt;li&gt;How did you clean them?&lt;/li&gt;
&lt;li&gt;How did you encode the data?&lt;/li&gt;
&lt;li&gt;Why did you choose the algorithm?&lt;/li&gt;
&lt;li&gt;How did you evaluate it?&lt;/li&gt;
&lt;li&gt;Does it generalize?&lt;/li&gt;
&lt;li&gt;What would you improve?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you can explain these, you are learning ML—not just memorizing scikit-learn commands.&lt;/p&gt;




&lt;h1&gt;
  
  
  102. A simple explanation of your five-day journey
&lt;/h1&gt;

&lt;p&gt;If someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What did you actually learn in five days?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A strong answer is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I learned that machine learning is not just about selecting an algorithm. I started by understanding datasets using Pandas, cleaning missing and incorrectly formatted values, and using visualization to understand distributions and relationships. Then I learned the difference between classification and regression, converted categorical data into numerical representations, split data into training and testing sets, trained several models using scikit-learn, and compared their performance on unseen data. Finally, I experimented with scaling, transformations, ensemble models, hyperparameters, and even built a simple Gradio interface around a churn model.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much stronger story than:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I learned Logistic Regression, SVM and Random Forest."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  103. Suggested article structure for dev.to
&lt;/h1&gt;

&lt;p&gt;For your final public article, I recommend this structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Title

1. Why I started learning ML
2. What I thought ML was vs what I learned
3. AI vs ML
4. Types of ML
5. My five-day learning roadmap
6. Day 1 — Understanding and cleaning data
7. EDA and why graphs matter
8. Day 2 — Classification with customer churn
9. Encoding categorical data
10. Train/test split
11. Logistic Regression, SVM and Decision Trees
12. How I evaluated classification models
13. Day 3 — Turning the model into an application
14. Day 4/5 — Regression and Customer Lifetime Value
15. Linear Regression vs KNN vs SVR vs Trees vs Ensembles
16. Log transformation
17. Scaling
18. Hyperparameters
19. What I got wrong / what I would improve
20. What I learned
21. What I plan to learn next
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This will make the article feel like a &lt;strong&gt;journey&lt;/strong&gt;, rather than a textbook.&lt;/p&gt;




&lt;h1&gt;
  
  
  104. Recommended article narrative
&lt;/h1&gt;

&lt;p&gt;The most interesting story is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"I started with almost no ML knowledge."

        ↓

"I learned that data preparation matters."

        ↓

"I discovered that graphs help me ask questions
before training models."

        ↓

"I learned classification."

        ↓

"I learned regression."

        ↓

"I tried different algorithms and realized
there is no single best model."

        ↓

"I learned that a high training score
doesn't necessarily mean a good model."

        ↓

"I turned one model into a small application."

        ↓

"I now understand the basic ML workflow
and know what I need to learn next."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a much more authentic beginner-to-ML story.&lt;/p&gt;




&lt;h1&gt;
  
  
  105. What to learn next
&lt;/h1&gt;

&lt;p&gt;Based on what your notebooks already cover, I would &lt;strong&gt;not&lt;/strong&gt; recommend immediately jumping into deep learning.&lt;/p&gt;

&lt;p&gt;First strengthen these foundations:&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Statistics
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;mean&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;variance&lt;/li&gt;
&lt;li&gt;standard deviation&lt;/li&gt;
&lt;li&gt;distributions&lt;/li&gt;
&lt;li&gt;probability&lt;/li&gt;
&lt;li&gt;correlation&lt;/li&gt;
&lt;li&gt;conditional probability&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. ML evaluation
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;confusion matrix&lt;/li&gt;
&lt;li&gt;precision&lt;/li&gt;
&lt;li&gt;recall&lt;/li&gt;
&lt;li&gt;F1&lt;/li&gt;
&lt;li&gt;ROC-AUC&lt;/li&gt;
&lt;li&gt;MAE&lt;/li&gt;
&lt;li&gt;RMSE&lt;/li&gt;
&lt;li&gt;R²&lt;/li&gt;
&lt;li&gt;cross-validation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Feature engineering
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;handling dates&lt;/li&gt;
&lt;li&gt;categorical features&lt;/li&gt;
&lt;li&gt;transformations&lt;/li&gt;
&lt;li&gt;interactions&lt;/li&gt;
&lt;li&gt;outliers&lt;/li&gt;
&lt;li&gt;missing values&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Model selection
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;cross-validation&lt;/li&gt;
&lt;li&gt;hyperparameter tuning&lt;/li&gt;
&lt;li&gt;GridSearchCV&lt;/li&gt;
&lt;li&gt;RandomizedSearchCV&lt;/li&gt;
&lt;li&gt;pipelines&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. ML fundamentals
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;bias vs variance&lt;/li&gt;
&lt;li&gt;underfitting&lt;/li&gt;
&lt;li&gt;overfitting&lt;/li&gt;
&lt;li&gt;regularization&lt;/li&gt;
&lt;li&gt;data leakage&lt;/li&gt;
&lt;li&gt;feature importance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then move toward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Classical ML
     ↓
Feature engineering
     ↓
Model selection
     ↓
Cross-validation
     ↓
ML pipelines
     ↓
Deployment
     ↓
Deep Learning
     ↓
Generative AI / LLMs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  106. The one-page cheat sheet
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Problem type
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Predict a category → Classification
Predict a number   → Regression
Find hidden groups → Clustering
Learn through reward → Reinforcement Learning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Data
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Rows       → observations
Columns    → variables/features
X          → input features
y          → target
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  First inspection
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Visualization
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Countplot   → categories
Histogram   → numerical distribution
Boxplot     → spread/outliers
Scatterplot → two numerical variables
Pairplot    → many numerical relationships
Heatmap     → correlation matrix
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Preparation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Missing values
Data types
Categorical encoding
Feature selection
Scaling when needed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Training
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Prediction
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Evaluation
&lt;/h2&gt;

&lt;p&gt;Classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accuracy
Precision
Recall
F1
Confusion Matrix
ROC-AUC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Regression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R²
MAE
MSE
RMSE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Core warning signs
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score &amp;gt;&amp;gt; Test score
        ↓
Possible overfitting

Test information used during preprocessing
        ↓
Possible data leakage

Identifier used as feature
        ↓
Possible meaningless pattern

Categorical values converted to arbitrary numbers
        ↓
Possible false ordering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  107. Final takeaway
&lt;/h1&gt;

&lt;p&gt;The biggest lesson from these five days is not a particular Python library or algorithm.&lt;/p&gt;

&lt;p&gt;It is the workflow:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Understand the problem → understand the data → clean the data → explore the data → prepare features → train → evaluate → improve → deploy.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The algorithms are tools inside that workflow.&lt;/p&gt;

&lt;p&gt;You have already touched many of the important building blocks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pandas&lt;/li&gt;
&lt;li&gt;NumPy&lt;/li&gt;
&lt;li&gt;Matplotlib&lt;/li&gt;
&lt;li&gt;Seaborn&lt;/li&gt;
&lt;li&gt;Scikit-learn&lt;/li&gt;
&lt;li&gt;data cleaning&lt;/li&gt;
&lt;li&gt;EDA&lt;/li&gt;
&lt;li&gt;classification&lt;/li&gt;
&lt;li&gt;regression&lt;/li&gt;
&lt;li&gt;encoding&lt;/li&gt;
&lt;li&gt;train/test split&lt;/li&gt;
&lt;li&gt;scaling&lt;/li&gt;
&lt;li&gt;model evaluation&lt;/li&gt;
&lt;li&gt;hyperparameters&lt;/li&gt;
&lt;li&gt;ensemble models&lt;/li&gt;
&lt;li&gt;target transformation&lt;/li&gt;
&lt;li&gt;Gradio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next step is not to memorize more algorithms.&lt;/p&gt;

&lt;p&gt;The next step is to become comfortable answering:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why am I doing this step?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you can answer that question for every important line in your notebook, you are moving from "following an ML tutorial" to actually understanding machine learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  Appendix A — Notebook-by-notebook step map
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;11_Aug_mpg.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Notebook step&lt;/th&gt;
&lt;th&gt;Why it was done&lt;/th&gt;
&lt;th&gt;ML concept&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Import NumPy/Pandas/Matplotlib/Seaborn&lt;/td&gt;
&lt;td&gt;Load tools&lt;/td&gt;
&lt;td&gt;Python ML stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;read_csv()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Load data&lt;/td&gt;
&lt;td&gt;Data ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nunique()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand uniqueness&lt;/td&gt;
&lt;td&gt;Data exploration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;head()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect examples&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;shape&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand dataset size&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;describe()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Summary statistics&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;info()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Types/non-null values&lt;/td&gt;
&lt;td&gt;Data quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;countplot()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect categorical distribution&lt;/td&gt;
&lt;td&gt;Visualization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hist()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect numerical distributions&lt;/td&gt;
&lt;td&gt;Visualization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;isnull()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Detect actual missing values&lt;/td&gt;
&lt;td&gt;Data cleaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pairplot()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Explore relationships&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sample()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect random examples&lt;/td&gt;
&lt;td&gt;Data quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identify &lt;code&gt;car name&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Recognize identifier-like field&lt;/td&gt;
&lt;td&gt;Feature selection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replace &lt;code&gt;?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Convert fake missing values&lt;/td&gt;
&lt;td&gt;Data cleaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;to_numeric()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Make horsepower numerical&lt;/td&gt;
&lt;td&gt;Data preparation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median imputation&lt;/td&gt;
&lt;td&gt;Fill missing horsepower&lt;/td&gt;
&lt;td&gt;Missing-value handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train/test split&lt;/td&gt;
&lt;td&gt;Separate learning/evaluation data&lt;/td&gt;
&lt;td&gt;Generalization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Logistic Regression&lt;/td&gt;
&lt;td&gt;Start classification experiment&lt;/td&gt;
&lt;td&gt;Supervised learning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Important caveat
&lt;/h3&gt;

&lt;p&gt;The notebook switches from Auto MPG to a &lt;code&gt;loan_prediction.csv&lt;/code&gt; dataset near the end. That appears to be a copied/reused modeling experiment rather than a continuation of the Auto MPG workflow. For a polished article, present these as separate experiments rather than one continuous Auto MPG pipeline.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;code&gt;12th-Aug-Telco-Customer-churn.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Load Telco CSV&lt;/td&gt;
&lt;td&gt;Get customer data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;head()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;shape&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand size&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;describe()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Statistical overview&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop &lt;code&gt;customerID&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Remove identifier from features&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;info()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Check data types&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sample()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect random rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot of Churn&lt;/td&gt;
&lt;td&gt;Check class distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot of gender&lt;/td&gt;
&lt;td&gt;Explore category distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Histograms&lt;/td&gt;
&lt;td&gt;Explore numerical distributions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pairplot&lt;/td&gt;
&lt;td&gt;Explore relationships&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Convert TotalCharges&lt;/td&gt;
&lt;td&gt;Fix numeric type&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fill missing TotalCharges&lt;/td&gt;
&lt;td&gt;Handle missing data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LabelEncoder&lt;/td&gt;
&lt;td&gt;Convert categories to numbers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train/test split&lt;/td&gt;
&lt;td&gt;Evaluate on unseen data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Logistic Regression&lt;/td&gt;
&lt;td&gt;Classification model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SVC&lt;/td&gt;
&lt;td&gt;Alternative classifier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision Tree&lt;/td&gt;
&lt;td&gt;Alternative classifier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gradio&lt;/td&gt;
&lt;td&gt;Turn model into interactive app&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Important caveat
&lt;/h3&gt;

&lt;p&gt;The notebook uses LabelEncoder on many categorical input columns. This is useful for learning encoding, but one-hot encoding is generally safer for nominal multi-category features.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;code&gt;12-Aug-2.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;This notebook is a more developed version of the Telco project.&lt;/p&gt;

&lt;p&gt;Important improvements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dropping &lt;code&gt;customerID&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;converting &lt;code&gt;TotalCharges&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;one-hot encoding multi-category variables&lt;/li&gt;
&lt;li&gt;using &lt;code&gt;random_state=42&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;using &lt;code&gt;stratify=Y&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;scaling features&lt;/li&gt;
&lt;li&gt;training Logistic Regression&lt;/li&gt;
&lt;li&gt;building a Gradio prediction interface&lt;/li&gt;
&lt;li&gt;preserving training feature columns&lt;/li&gt;
&lt;li&gt;applying the same preprocessing to new input&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the notebook that most clearly demonstrates the transition from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ML experiment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;small ML application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  &lt;code&gt;13-Aug.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Main steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load AutoInsurance
       ↓
Inspect dataset
       ↓
Drop Customer / date columns
       ↓
EDA
       ↓
Identify numerical/categorical columns
       ↓
Generate many visualizations
       ↓
Encode categorical variables
       ↓
Set CLV as target
       ↓
Train/test split
       ↓
Try multiple regressors
       ↓
Try log target
       ↓
Try scaling
       ↓
Tune tree/ensemble hyperparameters
       ↓
Generate predictions
       ↓
Save prediction CSVs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This notebook demonstrates the transition from classification to regression and from a single baseline model to model comparison and tuning.&lt;/p&gt;




&lt;h1&gt;
  
  
  Appendix B — A cleaner ML template to remember
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 1. Load
&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Understand
&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Clean
# handle missing values
# fix data types
# remove inappropriate identifiers
&lt;/span&gt;
&lt;span class="c1"&gt;# 4. Explore
# histograms
# boxplots
# countplots
# scatterplots
# correlation
&lt;/span&gt;
&lt;span class="c1"&gt;# 5. Prepare
&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# 6. Split
&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 7. Preprocess
# encoding / scaling
&lt;/span&gt;
&lt;span class="c1"&gt;# 8. Train
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 9. Predict
&lt;/span&gt;&lt;span class="n"&gt;predictions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 10. Evaluate
# choose metrics appropriate to the problem
&lt;/span&gt;
&lt;span class="c1"&gt;# 11. Compare / improve
# try another model
# tune hyperparameters
# use cross-validation
&lt;/span&gt;
&lt;span class="c1"&gt;# 12. Final model
# train using the selected approach
&lt;/span&gt;
&lt;span class="c1"&gt;# 13. Predict new data
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Closing thought for the article
&lt;/h1&gt;

&lt;p&gt;Five days ago, an ML notebook could easily look like a collection of unfamiliar commands.&lt;/p&gt;

&lt;p&gt;Now there is a structure behind those commands.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;head()&lt;/code&gt; is not just a command — it is a way to understand the data.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;isnull()&lt;/code&gt; is not just a command — it is part of data quality.&lt;/p&gt;

&lt;p&gt;A histogram is not just a graph — it tells you about a distribution.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;train_test_split()&lt;/code&gt; is not just boilerplate — it protects your evaluation from simply measuring memorization.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;fit()&lt;/code&gt; is not magic — it is the learning stage.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;predict()&lt;/code&gt; is the model applying what it learned.&lt;/p&gt;

&lt;p&gt;And model comparison is not about finding the fanciest algorithm — it is about finding an approach that generalizes well to data the model has never seen.&lt;/p&gt;

&lt;p&gt;That, more than anything else, is what I took away from my first five days of machine learning.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI &amp; ML Foundations: Cleaning Data, Building Regression &amp; Classification Models in Python</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Mon, 31 Aug 2026 16:49:07 +0000</pubDate>
      <link>https://dev.to/yashnigam/ai-ml-foundations-cleaning-data-building-regression-classification-models-in-python-n53</link>
      <guid>https://dev.to/yashnigam/ai-ml-foundations-cleaning-data-building-regression-classification-models-in-python-n53</guid>
      <description>&lt;h1&gt;
  
  
  My 5-Day Journey into Machine Learning: From Data Cleaning to My First ML Models
&lt;/h1&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Learning goal:&lt;/strong&gt; This is not a collection of code snippets to memorize. It is a beginner-friendly guide to understanding &lt;em&gt;why&lt;/em&gt; each step in a typical machine-learning workflow exists, what problem it solves, and how the pieces fit together.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A note about the source material
&lt;/h2&gt;

&lt;p&gt;This guide was built primarily from the four Jupyter notebooks you provided:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;11_Aug_mpg.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;12-Aug-2.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;12th-Aug-Telco-Customer-churn.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;13-Aug.ipynb&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;123&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your PDF notes were also provided, but the PDF is image-based and its text could not be reliably extracted. So I have &lt;strong&gt;not invented or silently reconstructed&lt;/strong&gt; material from the PDF. The explanations below are grounded mainly in the notebooks, with general ML knowledge added to explain the concepts behind the code.&lt;/p&gt;




&lt;h1&gt;
  
  
  1. Overview: What did I actually learn in these five days?
&lt;/h1&gt;

&lt;p&gt;At first glance, machine learning can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load dataset
    ↓
Clean data
    ↓
Create graphs
    ↓
Train model
    ↓
Check accuracy
    ↓
Done!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But that misses the most important part.&lt;/p&gt;

&lt;p&gt;The real ML workflow is closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Business / real-world question
        ↓
Define what we want to predict
        ↓
Understand the data
        ↓
Clean and prepare the data
        ↓
Explore the data visually
        ↓
Choose useful features
        ↓
Convert data into numbers
        ↓
Split into training and testing data
        ↓
Choose an appropriate ML algorithm
        ↓
Train the model
        ↓
Evaluate it on unseen data
        ↓
Compare models
        ↓
Improve preprocessing / model settings
        ↓
Use the model to make predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important insight is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Machine learning is not mainly about choosing an algorithm. It is about turning a real-world question and messy data into a reliable prediction process.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Your notebooks actually demonstrate most of this workflow.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. What is Artificial Intelligence?
&lt;/h1&gt;

&lt;p&gt;Artificial Intelligence (AI) is the broad idea of building systems that can perform tasks that normally require some form of human intelligence.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;recognizing an image&lt;/li&gt;
&lt;li&gt;understanding language&lt;/li&gt;
&lt;li&gt;recommending a movie&lt;/li&gt;
&lt;li&gt;detecting fraud&lt;/li&gt;
&lt;li&gt;predicting whether a customer will leave&lt;/li&gt;
&lt;li&gt;predicting the price/value of something&lt;/li&gt;
&lt;li&gt;generating text or images&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A useful mental model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Artificial Intelligence
│
├── Machine Learning
│   │
│   ├── Supervised Learning
│   ├── Unsupervised Learning
│   └── Reinforcement Learning
│
├── Deep Learning
│
└── Generative AI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Fyash-nigam%2FAI-ML-Foundations%2Fblob%2F7516a4b681369d4b105ca19706f7d06da65e30a9%2Fimages%2FAIMLHierarchy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2Fyash-nigam%2FAI-ML-Foundations%2Fblob%2F7516a4b681369d4b105ca19706f7d06da65e30a9%2Fimages%2FAIMLHierarchy.png" alt="image" width="" height=""&gt;&lt;/a&gt;&lt;br&gt;
This is simplified because these areas overlap, but it is a useful beginner mental model.&lt;/p&gt;


&lt;h1&gt;
  
  
  3. What is Machine Learning?
&lt;/h1&gt;

&lt;p&gt;Machine learning is a way of building systems where the computer learns patterns from data instead of us manually writing every rule.&lt;/p&gt;

&lt;p&gt;Imagine we want to predict whether a telecom customer will churn.&lt;/p&gt;

&lt;p&gt;A traditional rule-based approach might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF customer has a month-to-month contract
AND monthly charges are high
AND tenure is low
THEN churn = yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem is that we have to invent those rules ourselves.&lt;/p&gt;

&lt;p&gt;With machine learning, we give the algorithm historical examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer information → Actual churn result

Customer A → No
Customer B → Yes
Customer C → No
Customer D → Yes
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The algorithm tries to learn the relationship between the input information and the known outcome.&lt;/p&gt;

&lt;p&gt;Then we can give it a new customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New customer information
        ↓
      Model
        ↓
Predicted churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the core idea of supervised machine learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. The most important ML vocabulary
&lt;/h1&gt;

&lt;p&gt;Before learning algorithms, understand these words.&lt;/p&gt;

&lt;h2&gt;
  
  
  4.1 Dataset
&lt;/h2&gt;

&lt;p&gt;A dataset is a collection of examples.&lt;/p&gt;

&lt;p&gt;In your notebooks, a dataset is loaded into a Pandas DataFrame.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Telco_Customer_Churn.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Think of the DataFrame as a spreadsheet.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.2 Row / observation / sample
&lt;/h2&gt;

&lt;p&gt;Each row represents one example.&lt;/p&gt;

&lt;p&gt;For the Telco dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one customer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the Auto MPG dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one car
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the insurance dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one row = one customer/policy record
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These words are often used interchangeably:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;row&lt;/li&gt;
&lt;li&gt;observation&lt;/li&gt;
&lt;li&gt;sample&lt;/li&gt;
&lt;li&gt;instance&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4.3 Feature
&lt;/h2&gt;

&lt;p&gt;A feature is an input variable that the model can use to make a prediction.&lt;/p&gt;

&lt;p&gt;For customer churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tenure
MonthlyCharges
Contract
InternetService
PaymentMethod
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are features.&lt;/p&gt;

&lt;p&gt;A useful question to ask is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What information is available to the model before it makes its prediction?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That information is generally your feature set.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.4 Target
&lt;/h2&gt;

&lt;p&gt;The target is what we are trying to predict.&lt;/p&gt;

&lt;p&gt;Examples from your notebooks:&lt;/p&gt;

&lt;h3&gt;
  
  
  Telco
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Target = Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model predicts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0 → No churn
1 → Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Insurance
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Target = Customer Lifetime Value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model predicts a numerical value.&lt;/p&gt;




&lt;h2&gt;
  
  
  4.5 X and y
&lt;/h2&gt;

&lt;p&gt;You repeatedly used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the most important patterns in supervised ML.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X = inputs/features
Y = target/output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X
├── tenure
├── MonthlyCharges
├── Contract
├── InternetService
└── ...

Y
└── Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;X = what the model knows&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Y = what we want the model to learn to predict&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  5. The three major types of Machine Learning
&lt;/h1&gt;

&lt;h2&gt;
  
  
  5.1 Supervised Learning
&lt;/h2&gt;

&lt;p&gt;In supervised learning, we have examples where the correct answer is already known.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input X → Known answer Y
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model learns the relationship between X and Y.&lt;/p&gt;

&lt;p&gt;Your notebooks mainly use supervised learning.&lt;/p&gt;

&lt;p&gt;There are two major supervised learning tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Classification
&lt;/h3&gt;

&lt;p&gt;The target is a category.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Spam / Not Spam
Churn / No Churn
Fraud / Not Fraud
Disease / No Disease
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your Telco project is a classification problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Regression
&lt;/h3&gt;

&lt;p&gt;The target is a numerical value.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;House price
Temperature
Sales
Customer Lifetime Value
Fuel efficiency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your Customer Lifetime Value project is a regression problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. Classification vs Regression
&lt;/h1&gt;

&lt;p&gt;This distinction is extremely important.&lt;/p&gt;

&lt;h2&gt;
  
  
  Classification
&lt;/h2&gt;

&lt;p&gt;Question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which category does this example belong to?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer → Churn or No Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Possible output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0
1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common classification algorithms you tried:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Logistic Regression&lt;/li&gt;
&lt;li&gt;Support Vector Classifier (SVC)&lt;/li&gt;
&lt;li&gt;Decision Tree Classifier&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Regression
&lt;/h2&gt;

&lt;p&gt;Question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What numerical value should we predict?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer → Customer Lifetime Value = 6,543.21
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common regression algorithms you tried:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Linear Regression&lt;/li&gt;
&lt;li&gt;KNN Regressor&lt;/li&gt;
&lt;li&gt;SVR&lt;/li&gt;
&lt;li&gt;Decision Tree Regressor&lt;/li&gt;
&lt;li&gt;Bagging Regressor&lt;/li&gt;
&lt;li&gt;AdaBoost Regressor&lt;/li&gt;
&lt;li&gt;Random Forest Regressor&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  7. Unsupervised Learning
&lt;/h1&gt;

&lt;p&gt;In unsupervised learning, we do not have a known target.&lt;/p&gt;

&lt;p&gt;Instead, the algorithm tries to discover structure in the data.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer data
     ↓
Find groups
     ↓
Group 1: high-value customers
Group 2: price-sensitive customers
Group 3: new customers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;K-Means clustering&lt;/li&gt;
&lt;li&gt;Hierarchical clustering&lt;/li&gt;
&lt;li&gt;PCA&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You did not build an unsupervised-learning project in these notebooks, but it is important to know where it fits.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. Reinforcement Learning
&lt;/h1&gt;

&lt;p&gt;Reinforcement learning is different.&lt;/p&gt;

&lt;p&gt;An agent interacts with an environment and receives rewards or penalties.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓ action
Environment
  ↓
Reward / penalty
  ↓
Agent learns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;game-playing agents&lt;/li&gt;
&lt;li&gt;robotics&lt;/li&gt;
&lt;li&gt;certain recommendation/control systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This was not part of your five-day practical work.&lt;/p&gt;




&lt;h1&gt;
  
  
  9. Different ways to approach an ML problem
&lt;/h1&gt;

&lt;p&gt;A useful way to think about ML is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which algorithm should I use?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead ask:&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1 — What is the question?
&lt;/h3&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can I predict whether a telecom customer will churn?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Step 2 — What type of target do I have?
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn → category → Classification
Customer Lifetime Value → number → Regression
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 3 — What data is available?
&lt;/h3&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What columns exist?&lt;/li&gt;
&lt;li&gt;Which are numerical?&lt;/li&gt;
&lt;li&gt;Which are categorical?&lt;/li&gt;
&lt;li&gt;Are values missing?&lt;/li&gt;
&lt;li&gt;Are there strange values?&lt;/li&gt;
&lt;li&gt;Are some columns identifiers?&lt;/li&gt;
&lt;li&gt;Is the target balanced?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 4 — Explore before modeling
&lt;/h3&gt;

&lt;p&gt;Use statistics and graphs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5 — Prepare the data
&lt;/h3&gt;

&lt;p&gt;Convert it into a form the algorithm can understand.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6 — Establish a baseline
&lt;/h3&gt;

&lt;p&gt;Train a simple model first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 7 — Compare alternatives
&lt;/h3&gt;

&lt;p&gt;Try other appropriate models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 8 — Evaluate on unseen data
&lt;/h3&gt;

&lt;p&gt;A model is useful only if it works beyond the data it memorized.&lt;/p&gt;




&lt;h1&gt;
  
  
  10. Google Colab: Why did we use it?
&lt;/h1&gt;

&lt;p&gt;Google Colab gives you a cloud-based Jupyter Notebook environment.&lt;/p&gt;

&lt;p&gt;Instead of installing everything locally, you can run Python code in the browser.&lt;/p&gt;

&lt;p&gt;A notebook combines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Explanation
+
Python code
+
Output
+
Graphs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes it very useful for learning ML.&lt;/p&gt;

&lt;p&gt;Your notebooks use paths such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;auto&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;mpg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;csv&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is typical of a Colab environment.&lt;/p&gt;

&lt;p&gt;A normal learning workflow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Upload dataset
      ↓
Open notebook
      ↓
Run cells
      ↓
Inspect output
      ↓
Change code
      ↓
Run again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  11. The Python libraries you used
&lt;/h1&gt;

&lt;h2&gt;
  
  
  11.1 NumPy
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;NumPy provides numerical operations and arrays.&lt;/p&gt;

&lt;p&gt;You used it for things such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;
&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expm1&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;code&gt;np.nan&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Represents a missing numerical value.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;np.log1p(x)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Calculates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;log(1 + x)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is useful when a target is heavily skewed.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;np.expm1(x)&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Reverses &lt;code&gt;log1p&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exp(x) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You used this later to convert log predictions back to the original CLV scale.&lt;/p&gt;




&lt;h1&gt;
  
  
  12. Pandas
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pandas is one of the most important Python libraries for data work.&lt;/p&gt;

&lt;p&gt;Think of a Pandas DataFrame as a spreadsheet that Python can manipulate.&lt;/p&gt;

&lt;p&gt;You used it to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;load CSV files&lt;/li&gt;
&lt;li&gt;inspect data&lt;/li&gt;
&lt;li&gt;remove columns&lt;/li&gt;
&lt;li&gt;convert data types&lt;/li&gt;
&lt;li&gt;detect missing values&lt;/li&gt;
&lt;li&gt;fill missing values&lt;/li&gt;
&lt;li&gt;select X and Y&lt;/li&gt;
&lt;li&gt;create encoded columns&lt;/li&gt;
&lt;li&gt;save predictions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Telco_Customer_Churn.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  13. Matplotlib
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;matplotlib.pyplot&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;plt&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Matplotlib is a general-purpose plotting library.&lt;/p&gt;

&lt;p&gt;It provides the foundation for many Python visualizations.&lt;/p&gt;

&lt;p&gt;You used it for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;figure&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;title&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;plt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;show&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  14. Seaborn
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;seaborn&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sns&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seaborn makes statistical visualizations easier to create.&lt;/p&gt;

&lt;p&gt;You used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;countplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;boxplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;histplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pairplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;scatterplot&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;heatmap&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A key lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A graph is not decoration. It is a tool for asking questions about the data.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  15. Scikit-learn
&lt;/h1&gt;

&lt;p&gt;Scikit-learn is the main ML library used in your notebooks.&lt;/p&gt;

&lt;p&gt;You used it for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;splitting data&lt;/li&gt;
&lt;li&gt;preprocessing&lt;/li&gt;
&lt;li&gt;encoding&lt;/li&gt;
&lt;li&gt;scaling&lt;/li&gt;
&lt;li&gt;classification&lt;/li&gt;
&lt;li&gt;regression&lt;/li&gt;
&lt;li&gt;evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The basic pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SomeModel&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;predictions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This pattern appears again and again in ML.&lt;/p&gt;




&lt;h1&gt;
  
  
  16. The most important workflow: EDA
&lt;/h1&gt;

&lt;p&gt;EDA means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Exploratory Data Analysis&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;EDA is the process of investigating the dataset before building the final model.&lt;/p&gt;

&lt;p&gt;Your notebooks use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;plus graphs.&lt;/p&gt;

&lt;p&gt;These are not random commands.&lt;/p&gt;

&lt;p&gt;They answer different questions.&lt;/p&gt;




&lt;h1&gt;
  
  
  17. &lt;code&gt;df.head()&lt;/code&gt; — What does the data look like?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It displays the first few rows.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because before doing anything else, you want to see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;column names&lt;/li&gt;
&lt;li&gt;example values&lt;/li&gt;
&lt;li&gt;obvious data problems&lt;/li&gt;
&lt;li&gt;whether the data loaded correctly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Think of it as opening the box before using what is inside.&lt;/p&gt;




&lt;h1&gt;
  
  
  18. &lt;code&gt;df.shape&lt;/code&gt; — How much data do I have?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Auto MPG:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(398, 9)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;398 rows
9 columns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why does this matter?&lt;/p&gt;

&lt;p&gt;Because dataset size affects:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model choice&lt;/li&gt;
&lt;li&gt;computation time&lt;/li&gt;
&lt;li&gt;confidence in results&lt;/li&gt;
&lt;li&gt;risk of overfitting&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  19. &lt;code&gt;df.info()&lt;/code&gt; — What types of data do I have?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This tells you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;column names&lt;/li&gt;
&lt;li&gt;number of non-null values&lt;/li&gt;
&lt;li&gt;data types&lt;/li&gt;
&lt;li&gt;memory usage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is extremely important.&lt;/p&gt;

&lt;p&gt;For example, in Auto MPG:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;horsepower → object
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance horsepower sounds numerical.&lt;/p&gt;

&lt;p&gt;But Pandas sees it as text.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because the dataset contains values such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A column containing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;130
165
150
?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;cannot be treated as a clean numerical column.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;df.info()&lt;/code&gt; helps you detect data-type problems before modeling.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  20. &lt;code&gt;df.describe()&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and sometimes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives summary statistics.&lt;/p&gt;

&lt;p&gt;For numerical data you get things such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;count&lt;/li&gt;
&lt;li&gt;mean&lt;/li&gt;
&lt;li&gt;standard deviation&lt;/li&gt;
&lt;li&gt;minimum&lt;/li&gt;
&lt;li&gt;25th percentile&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;75th percentile&lt;/li&gt;
&lt;li&gt;maximum&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This helps you understand the scale and distribution of variables.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;minimum
   ↓
25%
   ↓
median
   ↓
75%
   ↓
maximum
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A large difference between the 75th percentile and maximum can be a clue that there may be extreme values.&lt;/p&gt;




&lt;h1&gt;
  
  
  21. &lt;code&gt;df.sample()&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or similar.&lt;/p&gt;

&lt;p&gt;This shows random rows.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because the first rows may not represent the whole dataset.&lt;/p&gt;

&lt;p&gt;Random samples can expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;unusual values&lt;/li&gt;
&lt;li&gt;formatting issues&lt;/li&gt;
&lt;li&gt;unexpected categories&lt;/li&gt;
&lt;li&gt;data-entry problems&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  22. Missing data
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many missing values are there in each column?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Missing data is important because many ML algorithms cannot directly work with missing values.&lt;/p&gt;

&lt;p&gt;But there is a subtle lesson from your Auto MPG notebook.&lt;/p&gt;

&lt;p&gt;You first checked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while &lt;code&gt;horsepower&lt;/code&gt; still contained &lt;code&gt;"?"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Pandas did not count &lt;code&gt;"?"&lt;/code&gt; as a missing value because &lt;code&gt;"?"&lt;/code&gt; is a string, not a true &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"?" ≠ NaN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So data cleaning sometimes requires identifying &lt;strong&gt;fake missing values&lt;/strong&gt; first.&lt;/p&gt;




&lt;h1&gt;
  
  
  23. Auto MPG: cleaning horsepower
&lt;/h1&gt;

&lt;p&gt;Your notebook did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This converts the placeholder into a real missing value.&lt;/p&gt;

&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now Pandas can treat horsepower as numerical data.&lt;/p&gt;

&lt;p&gt;This is a very important data-cleaning pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;messy text
   ↓
identify invalid placeholder
   ↓
convert to NaN
   ↓
convert column to numeric
   ↓
handle missing values
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  24. Why use the median?
&lt;/h1&gt;

&lt;p&gt;You calculated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;median_horsepower&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and then filled missing values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;median_horsepower&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why median?&lt;/p&gt;

&lt;p&gt;Because the median is less affected by extreme values than the mean.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10, 11, 12, 13, 1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mean is pulled heavily upward by 1000.&lt;/p&gt;

&lt;p&gt;The median is much more representative of the middle of the data.&lt;/p&gt;

&lt;p&gt;A useful rule of thumb:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fairly symmetric data → mean may work well&lt;/li&gt;
&lt;li&gt;skewed data / outliers → median is often safer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;There is no universal rule that "median is always correct."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The best treatment depends on the data and the problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  25. Dropping an identifier
&lt;/h1&gt;

&lt;p&gt;You considered dropping:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;car name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and actually dropped &lt;code&gt;customerID&lt;/code&gt; in the Telco work.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;An identifier is usually not a meaningful predictive feature.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID = 759832
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not inherently mean the customer is more or less likely to churn.&lt;/p&gt;

&lt;p&gt;The ID exists to identify the row, not to describe the customer.&lt;/p&gt;

&lt;p&gt;Important distinction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A column can be useful for identifying a record without being useful for predicting the target.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  26. But don't blindly drop columns
&lt;/h1&gt;

&lt;p&gt;This is an important improvement to your original reasoning.&lt;/p&gt;

&lt;p&gt;A column should not be dropped merely because it is an object/string.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract
InternetService
PaymentMethod
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;are categorical strings, but they can contain very useful predictive information.&lt;/p&gt;

&lt;p&gt;So the correct question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does this column contain useful information that is available at prediction time?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is this column numeric?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  27. Understanding your graphs
&lt;/h1&gt;

&lt;p&gt;You created many graphs. The important thing is to understand what question each graph answers.&lt;/p&gt;




&lt;h1&gt;
  
  
  28. Countplot
&lt;/h1&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A countplot answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many observations are in each category?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No churn → number of customers
Churn    → number of customers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This immediately helps you understand class balance.&lt;/p&gt;

&lt;p&gt;If the graph looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No Churn ███████████████████
Churn    ███████
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the classes are imbalanced.&lt;/p&gt;

&lt;p&gt;That matters because a model could achieve high accuracy simply by favoring the majority class.&lt;/p&gt;




&lt;h1&gt;
  
  
  29. Countplot with &lt;code&gt;hue&lt;/code&gt;
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Contract&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;hue&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How does the target category vary across another category?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract type
     ↓
Churn / No churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If month-to-month customers show much more churn than two-year customers, that is an important pattern worth investigating.&lt;/p&gt;

&lt;p&gt;But remember:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A graph showing association does not automatically prove causation.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  30. Histogram
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hist&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;figsize&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;histplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MonthlyCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;kde&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A histogram shows the distribution of a numerical variable.&lt;/p&gt;

&lt;p&gt;It helps answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where are most values?&lt;/li&gt;
&lt;li&gt;Is the data symmetric?&lt;/li&gt;
&lt;li&gt;Is it skewed?&lt;/li&gt;
&lt;li&gt;Are there multiple groups?&lt;/li&gt;
&lt;li&gt;Are there extreme values?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;frequency
   ^
   |       ███
   |     ███████
   |   █████████
   | ███████████
   +-----------------&amp;gt; value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  31. Why distributions matter
&lt;/h1&gt;

&lt;p&gt;Suppose a feature looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Most values: 0–100
A few values: 10,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That may indicate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;genuine outliers&lt;/li&gt;
&lt;li&gt;a different population&lt;/li&gt;
&lt;li&gt;data-entry errors&lt;/li&gt;
&lt;li&gt;a highly skewed distribution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A model may react differently depending on the algorithm.&lt;/p&gt;

&lt;p&gt;This is one reason EDA happens before modeling.&lt;/p&gt;




&lt;h1&gt;
  
  
  32. Boxplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A boxplot summarizes a distribution.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;       outlier
          •
          |
      ┌───────┐
      │       │
------│  box  │------
      │       │
      └───────┘
          |
       outlier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The box represents the middle portion of the data.&lt;/p&gt;

&lt;p&gt;The line inside the box is the median.&lt;/p&gt;

&lt;p&gt;Points outside the whiskers may be treated as potential outliers.&lt;/p&gt;

&lt;p&gt;Important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An outlier is not automatically an error.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A high-income customer may be perfectly legitimate.&lt;/p&gt;




&lt;h1&gt;
  
  
  33. Boxplot: target vs category
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Coverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a powerful graph.&lt;/p&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How does the distribution of Customer Lifetime Value differ between categories?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Coverage A → CLV distribution
Coverage B → CLV distribution
Coverage C → CLV distribution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;spread&lt;/li&gt;
&lt;li&gt;outliers&lt;/li&gt;
&lt;li&gt;overlap&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This can help you identify potentially useful relationships.&lt;/p&gt;




&lt;h1&gt;
  
  
  34. Scatterplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A scatterplot examines the relationship between two numerical variables.&lt;/p&gt;

&lt;p&gt;Each dot represents an observation.&lt;/p&gt;

&lt;p&gt;It helps answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"As X changes, does Y appear to change?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Possible patterns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Positive relationship:
  •
    •
      •
        •

Negative relationship:
        •
      •
    •
  •

No obvious relationship:
 •    •
    •
  •      •
     •
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pattern does not have to be a straight line.&lt;/p&gt;




&lt;h1&gt;
  
  
  35. Pairplot
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pairplot&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A pairplot shows many numerical relationships at once.&lt;/p&gt;

&lt;p&gt;It is useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;spotting correlations&lt;/li&gt;
&lt;li&gt;identifying clusters&lt;/li&gt;
&lt;li&gt;seeing distributions&lt;/li&gt;
&lt;li&gt;finding obvious relationships&lt;/li&gt;
&lt;li&gt;detecting possible outliers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The downside is that it becomes difficult to read when there are many columns.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Pairplot is excellent for small-to-medium exploratory datasets, but not something you blindly run on every large dataset.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  36. Correlation heatmap
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;heatmap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MonthlyCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]].&lt;/span&gt;&lt;span class="nf"&gt;corr&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;annot&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Correlation measures how two numerical variables move together in a linear relationship.&lt;/p&gt;

&lt;p&gt;A correlation close to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+1 → strong positive linear relationship
 0 → little/no linear relationship
-1 → strong negative linear relationship
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tenure ↑
TotalCharges ↑
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;could result in positive correlation.&lt;/p&gt;

&lt;p&gt;But correlation has an important limitation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correlation does not prove causation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Also, zero correlation does not necessarily mean "no relationship" because the relationship might be non-linear.&lt;/p&gt;




&lt;h1&gt;
  
  
  37. A key lesson from visualization
&lt;/h1&gt;

&lt;p&gt;The graphs answer different questions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Graph&lt;/th&gt;
&lt;th&gt;Main question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Countplot&lt;/td&gt;
&lt;td&gt;How many observations are in each category?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot + hue&lt;/td&gt;
&lt;td&gt;How does one category vary with another?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Histogram&lt;/td&gt;
&lt;td&gt;What does a numerical distribution look like?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boxplot&lt;/td&gt;
&lt;td&gt;What are the median, spread and potential outliers?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boxplot + category&lt;/td&gt;
&lt;td&gt;How does a numerical distribution differ by category?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scatterplot&lt;/td&gt;
&lt;td&gt;How do two numerical variables relate?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pairplot&lt;/td&gt;
&lt;td&gt;What relationships exist among several numerical variables?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heatmap&lt;/td&gt;
&lt;td&gt;How strongly are numerical variables linearly correlated?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The goal is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Create lots of graphs."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The goal is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Use the right graph to answer the right question."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  38. Encoding: why does ML need it?
&lt;/h1&gt;

&lt;p&gt;Many ML algorithms work with numbers.&lt;/p&gt;

&lt;p&gt;But real-world data contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Male
Female

Yes
No

Month-to-month
One year
Two year

Fiber optic
DSL
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model needs these values represented numerically.&lt;/p&gt;

&lt;p&gt;This is called &lt;strong&gt;encoding&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  39. Label Encoding
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.preprocessing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LabelEncoder&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and encoded categorical columns.&lt;/p&gt;

&lt;p&gt;For a binary column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can be reasonable for a binary variable.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn:
No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes intuitive sense.&lt;/p&gt;




&lt;h1&gt;
  
  
  40. Why LabelEncoder can be dangerous for multi-category variables
&lt;/h1&gt;

&lt;p&gt;Suppose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Contract:
Month-to-month
One year
Two year
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Label encoding might produce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Month-to-month → 0
One year       → 1
Two year       → 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A numerical model may interpret that as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Two year &amp;gt; One year &amp;gt; Month-to-month
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That ordering may not be appropriate for the algorithm.&lt;/p&gt;

&lt;p&gt;The numbers are just codes.&lt;/p&gt;

&lt;p&gt;They do not automatically mean that category 2 is "twice" category 1.&lt;/p&gt;

&lt;p&gt;For nominal categories, one-hot encoding is usually safer.&lt;/p&gt;




&lt;h1&gt;
  
  
  41. One-hot encoding
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_dummies&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;drop_first&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This transforms categories into separate binary columns.&lt;/p&gt;

&lt;p&gt;For:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;InternetService:
DSL
Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we might create columns such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;InternetService_Fiber optic
InternetService_No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with the remaining category represented by both values being 0.&lt;/p&gt;

&lt;p&gt;This avoids inventing a fake numerical ordering.&lt;/p&gt;




&lt;h1&gt;
  
  
  42. Why &lt;code&gt;drop_first=True&lt;/code&gt;?
&lt;/h1&gt;

&lt;p&gt;If a categorical variable has several categories, using all one-hot columns can introduce redundancy in some models.&lt;/p&gt;

&lt;p&gt;Dropping one category creates a reference category.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DSL
Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;could become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fiber optic
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fiber optic = 0
No = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the reference category:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DSL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For beginner understanding, the key idea is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One-hot encoding converts categories into machine-readable indicator variables without pretending the categories have numerical order.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  43. Train/test split
&lt;/h1&gt;

&lt;p&gt;You repeatedly used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the most important ML concepts.&lt;/p&gt;

&lt;p&gt;Suppose we have 1,000 examples.&lt;/p&gt;

&lt;p&gt;We might use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;700 → training
300 → testing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The training set is used to learn the model.&lt;/p&gt;

&lt;p&gt;The test set is held back to evaluate how the trained model performs on unseen data.&lt;/p&gt;

&lt;p&gt;Think of it like an exam:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data = practice questions
Test data     = unseen exam questions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  44. Why not train and test on the same data?
&lt;/h1&gt;

&lt;p&gt;Because the model could simply memorize the training examples.&lt;/p&gt;

&lt;p&gt;Suppose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 99%
Test score     = 65%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a warning sign.&lt;/p&gt;

&lt;p&gt;The model learned the training data extremely well but does not generalize well.&lt;/p&gt;

&lt;p&gt;This is called:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Overfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  45. Underfitting
&lt;/h1&gt;

&lt;p&gt;The opposite can happen.&lt;/p&gt;

&lt;p&gt;If the model is too simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 60%
Test score     = 58%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It may not have learned enough from the data.&lt;/p&gt;

&lt;p&gt;This is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Underfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A useful mental picture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Underfitting
Model too simple
       ↓
misses important patterns

Good fit
Learns useful patterns
       ↓
works on unseen data

Overfitting
Model learns noise/details
       ↓
great training performance
poor unseen performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  46. Why &lt;code&gt;random_state&lt;/code&gt; matters
&lt;/h1&gt;

&lt;p&gt;In one Telco notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the improved notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A random state makes the split reproducible.&lt;/p&gt;

&lt;p&gt;Without it, the random split can change between runs.&lt;/p&gt;

&lt;p&gt;That means your score may change.&lt;/p&gt;

&lt;p&gt;With:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you can reproduce the same split.&lt;/p&gt;

&lt;p&gt;The number &lt;code&gt;42&lt;/code&gt; is not magical.&lt;/p&gt;

&lt;p&gt;Any fixed integer can serve this purpose.&lt;/p&gt;




&lt;h1&gt;
  
  
  47. Why &lt;code&gt;stratify=Y&lt;/code&gt; is useful for classification
&lt;/h1&gt;

&lt;p&gt;In your improved Telco notebook you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;stratify&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This helps preserve the class distribution between training and testing data.&lt;/p&gt;

&lt;p&gt;For example, if the full dataset contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;73% No Churn
27% Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;stratification aims to keep roughly the same proportion in both splits.&lt;/p&gt;

&lt;p&gt;This is particularly useful when classes are imbalanced.&lt;/p&gt;




&lt;h1&gt;
  
  
  48. Logistic Regression
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.linear_model&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LogisticRegression&lt;/span&gt;

&lt;span class="n"&gt;model_lr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LogisticRegression&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;model_lr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Despite its name, Logistic Regression is commonly used for &lt;strong&gt;classification&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Its goal is to estimate the probability of belonging to a class.&lt;/p&gt;

&lt;p&gt;For binary classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;features
   ↓
Logistic Regression
   ↓
probability
   ↓
class
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Probability of churn = 0.82
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model may classify that customer as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;depending on its decision threshold.&lt;/p&gt;




&lt;h1&gt;
  
  
  49. Why Logistic Regression is a good beginner model
&lt;/h1&gt;

&lt;p&gt;It is useful because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;relatively simple&lt;/li&gt;
&lt;li&gt;fast&lt;/li&gt;
&lt;li&gt;often strong as a baseline&lt;/li&gt;
&lt;li&gt;easier to interpret than many complex models&lt;/li&gt;
&lt;li&gt;naturally suited to binary classification&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A good ML habit is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Start with a simple baseline before reaching for complex models.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  50. Support Vector Machine / SVC
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.svm&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SVC&lt;/span&gt;

&lt;span class="n"&gt;model_lr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SVC&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The variable name &lt;code&gt;model_lr&lt;/code&gt; is misleading here.&lt;/p&gt;

&lt;p&gt;It is actually an SVC model.&lt;/p&gt;

&lt;p&gt;SVC means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Support Vector Classifier&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The basic idea is to find a decision boundary that separates classes.&lt;/p&gt;

&lt;p&gt;In simple 2D data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Class A: ● ● ●

--------- decision boundary ---------

Class B: ▲ ▲ ▲
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SVM tries to find a boundary with a useful margin between classes.&lt;/p&gt;




&lt;h1&gt;
  
  
  51. Why scaling matters for SVM
&lt;/h1&gt;

&lt;p&gt;SVM is sensitive to feature scale.&lt;/p&gt;

&lt;p&gt;Suppose one feature ranges from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0–1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and another ranges from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0–100,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The larger-scale feature can dominate distance-related calculations.&lt;/p&gt;

&lt;p&gt;That is why scaling is often important for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SVM&lt;/li&gt;
&lt;li&gt;KNN&lt;/li&gt;
&lt;li&gt;Logistic Regression in many situations&lt;/li&gt;
&lt;li&gt;neural networks&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  52. Decision Tree
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeClassifier&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A decision tree makes predictions using a sequence of questions.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is tenure &amp;lt; 12?
       /       \
     Yes       No
     /           \
Is charge &amp;gt; X?   ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eventually the tree reaches a leaf containing a prediction.&lt;/p&gt;

&lt;p&gt;This is attractive because it is easy to visualize conceptually.&lt;/p&gt;




&lt;h1&gt;
  
  
  53. Why trees can overfit
&lt;/h1&gt;

&lt;p&gt;A tree can keep splitting the data until it creates extremely specific rules.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If feature A &amp;gt; 10
and feature B &amp;lt; 4
and feature C = 7
and feature D &amp;gt; 92
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eventually the tree may memorize training examples.&lt;/p&gt;

&lt;p&gt;That is why you later tried:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;
&lt;span class="n"&gt;min_samples_split&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;
&lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are ways of controlling tree complexity.&lt;/p&gt;




&lt;h1&gt;
  
  
  54. Random Forest
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;RandomForestRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A random forest combines many decision trees.&lt;/p&gt;

&lt;p&gt;Instead of relying on one tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tree 1 ─┐
Tree 2 ─┤
Tree 3 ─┤
Tree 4 ─┤ → combined prediction
...     ┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The idea is that many different trees can produce a more robust prediction than a single tree.&lt;/p&gt;

&lt;p&gt;This is an example of an &lt;strong&gt;ensemble method&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  55. Bagging
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;BaggingRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_samples&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Bagging means combining models trained on different samples of the data.&lt;/p&gt;

&lt;p&gt;The overall idea:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dataset
  ↓
Different samples
  ↓
Multiple models
  ↓
Combine predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is often to reduce variance and make predictions more stable.&lt;/p&gt;

&lt;p&gt;Random Forest can be viewed as a specialized tree-based ensemble that adds additional randomness in how trees are built.&lt;/p&gt;




&lt;h1&gt;
  
  
  56. AdaBoost
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;AdaBoostRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AdaBoost means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Adaptive Boosting&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of training many independent models and simply averaging them, boosting builds models sequentially.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model 1
  ↓
find mistakes
  ↓
Model 2 focuses more on difficult cases
  ↓
Model 3 focuses further
  ↓
combine models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key idea:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Later models try to improve on the weaknesses of earlier models.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  57. K-Nearest Neighbors (KNN)
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;KNeighborsRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;KNN predicts using nearby examples.&lt;/p&gt;

&lt;p&gt;Imagine a new customer.&lt;/p&gt;

&lt;p&gt;The algorithm asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Which existing customers are most similar to this customer?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then it uses those neighbors to estimate the output.&lt;/p&gt;

&lt;p&gt;For regression, the prediction is commonly based on the neighbors' target values.&lt;/p&gt;




&lt;h1&gt;
  
  
  58. Why KNN needs scaling
&lt;/h1&gt;

&lt;p&gt;KNN relies on distance.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Income:   0–100,000
Age:      18–80
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If raw values are used, income can dominate the distance calculation.&lt;/p&gt;

&lt;p&gt;Scaling puts features onto comparable scales.&lt;/p&gt;

&lt;p&gt;That is why you experimented with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  59. StandardScaler
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;X_train_scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;X_test_scaled&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an important pattern.&lt;/p&gt;

&lt;p&gt;Notice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit_transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scaler&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;The scaler learns the training data's statistics.&lt;/p&gt;

&lt;p&gt;Then the exact same transformation is applied to the test data.&lt;/p&gt;

&lt;p&gt;You should not independently fit the scaler on the test set.&lt;/p&gt;

&lt;p&gt;Otherwise information from the test set leaks into preprocessing.&lt;/p&gt;

&lt;p&gt;This idea is called avoiding:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Data leakage&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  60. Data leakage
&lt;/h1&gt;

&lt;p&gt;Data leakage happens when information that should not be available to the model during training accidentally influences the model.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Train preprocessing
     ↓
uses information from test data
     ↓
test is no longer truly unseen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That can make evaluation look better than real-world performance.&lt;/p&gt;

&lt;p&gt;The correct principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fit preprocessing steps using training data, then apply them to validation/test/new data.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A production ML pipeline should ideally bundle preprocessing and modeling together.&lt;/p&gt;




&lt;h1&gt;
  
  
  61. Linear Regression
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;LinearRegression&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for Customer Lifetime Value.&lt;/p&gt;

&lt;p&gt;Linear regression tries to model a relationship like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = b0 + b1X1 + b2X2 + ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For one feature:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = intercept + slope × X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is to find coefficients that produce predictions close to the observed values.&lt;/p&gt;




&lt;h1&gt;
  
  
  62. Regression score: what does &lt;code&gt;.score()&lt;/code&gt; mean?
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the scikit-learn regression models you used, &lt;code&gt;.score()&lt;/code&gt; generally returns &lt;strong&gt;R²&lt;/strong&gt;, the coefficient of determination.&lt;/p&gt;

&lt;p&gt;Very roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R² = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the model explains the observed variation perfectly on that dataset.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R² = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;means the model is no better than a simple baseline based on predicting the mean target.&lt;/p&gt;

&lt;p&gt;R² can also be negative on unseen data.&lt;/p&gt;

&lt;p&gt;Important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;R² is not the same thing as accuracy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For regression, do not say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"My regression model has 90% accuracy"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;just because R² is 0.90.&lt;/p&gt;

&lt;p&gt;Instead say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The model achieved an R² of 0.90."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  63. Classification accuracy
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;accuracy_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test_pred&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Accuracy is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;correct predictions
-------------------
total predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;90 correct
100 total
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;gives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;90% accuracy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Accuracy is easy to understand.&lt;/p&gt;

&lt;p&gt;But it can be misleading when classes are highly imbalanced.&lt;/p&gt;




&lt;h1&gt;
  
  
  64. Confusion matrix
&lt;/h1&gt;

&lt;p&gt;Your later notes/workflow refer to concepts such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TP
TN
FP
FN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A confusion matrix organizes classification predictions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Actual
                 No       Yes
Predicted No     TN       FN
Predicted Yes    FP       TP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meaning:&lt;/p&gt;

&lt;h3&gt;
  
  
  True Positive
&lt;/h3&gt;

&lt;p&gt;Model predicted positive and it was positive.&lt;/p&gt;

&lt;h3&gt;
  
  
  True Negative
&lt;/h3&gt;

&lt;p&gt;Model predicted negative and it was negative.&lt;/p&gt;

&lt;h3&gt;
  
  
  False Positive
&lt;/h3&gt;

&lt;p&gt;Model predicted positive but it was negative.&lt;/p&gt;

&lt;h3&gt;
  
  
  False Negative
&lt;/h3&gt;

&lt;p&gt;Model predicted negative but it was positive.&lt;/p&gt;

&lt;p&gt;This is much more informative than accuracy alone when the cost of errors differs.&lt;/p&gt;




&lt;h1&gt;
  
  
  65. Precision and Recall
&lt;/h1&gt;

&lt;p&gt;For classification:&lt;/p&gt;

&lt;h3&gt;
  
  
  Precision
&lt;/h3&gt;

&lt;p&gt;Of everything the model predicted as positive:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many were actually positive?&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Precision = TP / (TP + FP)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Recall
&lt;/h3&gt;

&lt;p&gt;Of everything that was actually positive:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many did the model find?&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recall = TP / (TP + FN)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right metric depends on the problem.&lt;/p&gt;

&lt;p&gt;For churn prediction, for example, missing a customer who is about to churn may be more important than contacting an extra customer who would have stayed.&lt;/p&gt;




&lt;h1&gt;
  
  
  66. Your Telco Customer Churn project
&lt;/h1&gt;

&lt;p&gt;This is probably the clearest classification example in your notebooks.&lt;/p&gt;

&lt;p&gt;The business question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can we predict whether a telecom customer will churn?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The target is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Features are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Churn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer data
     ↓
Clean data
     ↓
Explore data
     ↓
Encode categories
     ↓
Split train/test
     ↓
Train classifier
     ↓
Predict churn
     ↓
Evaluate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  67. Telco: removing customerID
&lt;/h1&gt;

&lt;p&gt;You did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customerID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is reasonable because the ID is an identifier rather than a customer characteristic.&lt;/p&gt;

&lt;p&gt;You also commented that it could be kept separately if you later wanted to map predictions back to customers.&lt;/p&gt;

&lt;p&gt;That is a good practical idea.&lt;/p&gt;

&lt;p&gt;In production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID
   ↓
keep separately for business reporting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customerID
   X
do not necessarily use as a model feature
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  68. Telco: fixing TotalCharges
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TotalCharges&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coerce&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;errors="coerce"&lt;/code&gt; means invalid values are converted to &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Then you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;fillna&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in one notebook and median imputation in another.&lt;/p&gt;

&lt;p&gt;These are two different decisions.&lt;/p&gt;

&lt;p&gt;The important lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Missing-value handling should be based on the meaning of the data, not simply on whichever replacement is easiest.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example, if missing &lt;code&gt;TotalCharges&lt;/code&gt; occurs because a customer has zero tenure and has just joined, then filling with zero may have a meaningful business interpretation.&lt;/p&gt;

&lt;p&gt;If missingness represents an unknown measurement, median imputation may be more appropriate.&lt;/p&gt;




&lt;h1&gt;
  
  
  69. Telco: the first modeling version
&lt;/h1&gt;

&lt;p&gt;The earlier Telco notebook used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;LabelEncoder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;for many categorical columns.&lt;/p&gt;

&lt;p&gt;This is okay as a learning exercise for understanding encoding, but it is not ideal for every categorical variable.&lt;/p&gt;

&lt;p&gt;A better version in your later notebook used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_dummies&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;multi_cols&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;drop_first&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a meaningful improvement.&lt;/p&gt;




&lt;h1&gt;
  
  
  70. Telco: your improved preprocessing pipeline
&lt;/h1&gt;

&lt;p&gt;The later notebook did something closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw data
   ↓
Drop customerID
   ↓
Convert TotalCharges
   ↓
Encode binary columns
   ↓
One-hot encode multi-category columns
   ↓
Separate X and Y
   ↓
Scale X
   ↓
Train Logistic Regression
   ↓
Create Gradio interface
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is much closer to a real ML application.&lt;/p&gt;




&lt;h1&gt;
  
  
  71. Gradio: turning a model into an application
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;gradio&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;gr&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and built a UI.&lt;/p&gt;

&lt;p&gt;This is an important step because it changes the project from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Notebook experiment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User input
    ↓
Preprocessing
    ↓
Model
    ↓
Prediction
    ↓
Human-readable result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gender: Female
Tenure: 12 months
Contract: Month-to-month
Monthly charges: $70
...
          ↓
    Predict Churn
          ↓
High Churn Risk
Probability: ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This demonstrates an important ML engineering concept:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A model is useful only when it can be integrated into a workflow where people or systems can use its predictions.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  72. Very important: preprocessing must match training
&lt;/h1&gt;

&lt;p&gt;Your Gradio code contains a good lesson.&lt;/p&gt;

&lt;p&gt;During training you:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;encode the categorical values&lt;/li&gt;
&lt;li&gt;create dummy columns&lt;/li&gt;
&lt;li&gt;scale the features&lt;/li&gt;
&lt;li&gt;train the model&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;During prediction you must do the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;raw user input
    ↓
same encoding
    ↓
same columns
    ↓
same scaling
    ↓
model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the training data has:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature_1
feature_2
feature_3
feature_4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but the application sends:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature_1
feature_3
feature_4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model cannot interpret the input correctly.&lt;/p&gt;

&lt;p&gt;Your use of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;input_df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reindex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;columns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;feature_columns&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;fill_value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is intended to make the input columns match the training feature structure.&lt;/p&gt;

&lt;p&gt;That is an important practical idea.&lt;/p&gt;




&lt;h1&gt;
  
  
  73. A stronger production pattern: Pipeline
&lt;/h1&gt;

&lt;p&gt;A cleaner production approach is often:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;Pipeline&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preprocessor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;LogisticRegression&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then preprocessing and prediction are tied together.&lt;/p&gt;

&lt;p&gt;This reduces the risk of accidentally applying different transformations during training and inference.&lt;/p&gt;

&lt;p&gt;Your notebook demonstrates the concept manually, which is useful for learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  74. Your Auto Insurance / Customer Lifetime Value project
&lt;/h1&gt;

&lt;p&gt;The &lt;code&gt;13-Aug.ipynb&lt;/code&gt; notebook moves into regression.&lt;/p&gt;

&lt;p&gt;The target is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This changes the question from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Will this customer churn?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What numerical Customer Lifetime Value should we predict?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Therefore:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Classification → Churn
Regression     → Customer Lifetime Value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  75. Insurance dataset: initial exploration
&lt;/h1&gt;

&lt;p&gt;You loaded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AutoInsurance.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then inspected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You also removed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer
Effective To Date
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from the modeling DataFrame.&lt;/p&gt;

&lt;p&gt;Again, the purpose is to separate identifiers / dates that were not being used in the current modeling approach.&lt;/p&gt;

&lt;p&gt;However, an important ML lesson is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Dates are not automatically useless.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A date may contain useful information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;month&lt;/li&gt;
&lt;li&gt;day of week&lt;/li&gt;
&lt;li&gt;season&lt;/li&gt;
&lt;li&gt;year&lt;/li&gt;
&lt;li&gt;time since an event&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A stronger project could engineer useful date features instead of simply dropping every date column.&lt;/p&gt;




&lt;h1&gt;
  
  
  76. Insurance EDA
&lt;/h1&gt;

&lt;p&gt;You explored:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;countplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;State&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boxplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Coverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer Lifetime Value&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Income&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;hist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a good progression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Categorical distribution
        ↓
Numerical distribution
        ↓
Numerical vs categorical
        ↓
Numerical vs numerical
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You are gradually asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What might explain the target?"&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  77. Automatically finding numerical and categorical columns
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;numerical_cols&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_dtypes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;float64&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;

&lt;span class="n"&gt;categorical_cols&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select_dtypes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;columns&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful because instead of manually listing every column, Python can identify columns based on data type.&lt;/p&gt;

&lt;p&gt;Then you looped through them to create graphs.&lt;/p&gt;

&lt;p&gt;This is the beginning of writing reusable data-analysis code.&lt;/p&gt;




&lt;h1&gt;
  
  
  78. Automated boxplots
&lt;/h1&gt;

&lt;p&gt;You created boxplots for all numerical columns.&lt;/p&gt;

&lt;p&gt;This is useful because you can quickly inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;distributions&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;spread&lt;/li&gt;
&lt;li&gt;potential outliers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, remember that a large collection of graphs can become overwhelming.&lt;/p&gt;

&lt;p&gt;The goal should eventually move from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Let's plot everything."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Let's plot the variables that help answer our question."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a sign of growing from beginner EDA toward professional analysis.&lt;/p&gt;




&lt;h1&gt;
  
  
  79. Binning numerical variables for boxplots
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cut&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;col&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;bins&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This divides a numerical variable into five intervals.&lt;/p&gt;

&lt;p&gt;Then you compare CLV across those ranges.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Income range 1 → CLV distribution
Income range 2 → CLV distribution
Income range 3 → CLV distribution
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful for visualization because a boxplot expects categories on one axis.&lt;/p&gt;

&lt;p&gt;But the number and boundaries of bins can affect the story you see.&lt;/p&gt;

&lt;p&gt;So binning is mainly a visualization technique here, not necessarily something you should automatically use as a model feature.&lt;/p&gt;




&lt;h1&gt;
  
  
  80. Comparing many regression algorithms
&lt;/h1&gt;

&lt;p&gt;One of the biggest learning steps in your final notebook was trying multiple regressors:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Linear Regression
KNN Regressor
SVR
Decision Tree Regressor
Bagging Regressor
AdaBoost Regressor
Random Forest Regressor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is valuable because it teaches a core ML lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Different algorithms make different assumptions and capture different types of patterns.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is no universally best algorithm.&lt;/p&gt;




&lt;h1&gt;
  
  
  81. Linear Regression vs tree-based models
&lt;/h1&gt;

&lt;p&gt;Linear Regression assumes a relationship that can be represented through a linear combination of features.&lt;/p&gt;

&lt;p&gt;Tree-based methods can represent more complex non-linear relationships.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Linear model:

Y
│       /
│     /
│   /
│ /
└──────── X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A tree-based model can approximate more irregular patterns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y
│    ┌────
│    │
│ ───┘
│
└──────── X
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why trying multiple model families can be useful.&lt;/p&gt;




&lt;h1&gt;
  
  
  82. Why model comparison matters
&lt;/h1&gt;

&lt;p&gt;Suppose you obtain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model             Train R²    Test R²
Linear Regression   0.70       0.68
KNN                 0.90       0.62
Decision Tree       0.99       0.55
Random Forest       0.91       0.78
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should not automatically choose the model with the highest training score.&lt;/p&gt;

&lt;p&gt;The more important question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which model generalizes well to unseen data?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here Random Forest would be more interesting than the tree with 0.99 training R².&lt;/p&gt;




&lt;h1&gt;
  
  
  83. Log transformation of Customer Lifetime Value
&lt;/h1&gt;

&lt;p&gt;You later used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because some target variables are heavily skewed.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Most CLV values: relatively small
A few CLV values: extremely large
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model may have difficulty fitting such a distribution.&lt;/p&gt;

&lt;p&gt;A log transformation compresses large values.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Original:
1
10
100
1000
10000

Log scale:
small differences between large values
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can make the target distribution easier for some models to learn.&lt;/p&gt;




&lt;h1&gt;
  
  
  84. Reversing the log transformation
&lt;/h1&gt;

&lt;p&gt;After predicting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_actual&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expm1&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;If:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y_log = log(1 + Y)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Y = exp(Y_log) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;np.expm1()&lt;/code&gt; performs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exp(x) - 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;log target
   ↓
model
   ↓
log prediction
   ↓
expm1
   ↓
original target scale
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  85. An important issue in the notebook: split consistency
&lt;/h1&gt;

&lt;p&gt;In your log-transform section, you did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Y_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;X_train_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_train_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y_test_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;Y_log&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is okay as a new split, but it means your earlier train/test split and later train/test split are not necessarily the same.&lt;/p&gt;

&lt;p&gt;For a clean experiment, define the split once and reuse it consistently.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;y_train_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;y_test_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log1p&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes comparisons easier to reason about.&lt;/p&gt;




&lt;h1&gt;
  
  
  86. Another important issue: predicting on the full dataset
&lt;/h1&gt;

&lt;p&gt;You later did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;predictions_log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model_rf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and saved predictions for the full dataset.&lt;/p&gt;

&lt;p&gt;That can be useful for producing a business-facing prediction file.&lt;/p&gt;

&lt;p&gt;But these are &lt;strong&gt;not unbiased test predictions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because &lt;code&gt;X&lt;/code&gt; includes training observations the model already saw.&lt;/p&gt;

&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Predictions on X_test
→ useful for evaluating unseen data

Predictions on X
→ useful for generating predictions for all records,
   but not for measuring generalization
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction is very important.&lt;/p&gt;




&lt;h1&gt;
  
  
  87. Hyperparameters
&lt;/h1&gt;

&lt;p&gt;Later you changed models from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;DecisionTreeRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_split&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These settings are called &lt;strong&gt;hyperparameters&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;They are chosen by us before training.&lt;/p&gt;

&lt;p&gt;The model learns its internal parameters from the training data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Parameters
&lt;/h3&gt;

&lt;p&gt;Learned by the algorithm.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hyperparameters
&lt;/h3&gt;

&lt;p&gt;Set by us.&lt;/p&gt;

&lt;p&gt;This distinction is fundamental.&lt;/p&gt;




&lt;h1&gt;
  
  
  88. Understanding your tree hyperparameters
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;max_depth&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Controls how deep the tree can grow.&lt;/p&gt;

&lt;p&gt;Smaller:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;simpler tree
less overfitting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Larger:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;more complex tree
greater overfitting risk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  &lt;code&gt;min_samples_split&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Minimum number of samples required before a node can be split.&lt;/p&gt;

&lt;p&gt;Higher values make the tree more conservative.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;min_samples_leaf&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Minimum number of samples allowed in a leaf.&lt;/p&gt;

&lt;p&gt;Higher values prevent extremely tiny leaves.&lt;/p&gt;




&lt;h1&gt;
  
  
  89. Random Forest hyperparameters
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;RandomForestRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;min_samples_leaf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  &lt;code&gt;n_estimators=200&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Build 200 trees.&lt;/p&gt;

&lt;p&gt;More trees can make the ensemble more stable, though computation increases.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;max_depth=8&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Limits tree complexity.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;min_samples_leaf=5&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Prevents leaves from becoming too specific.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;random_state=42&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Makes the experiment reproducible.&lt;/p&gt;




&lt;h1&gt;
  
  
  90. Bagging hyperparameters
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;BaggingRegressor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;n_estimators&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_samples&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_features&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This means, conceptually:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;build many estimators&lt;/li&gt;
&lt;li&gt;each estimator sees a sample of the training rows&lt;/li&gt;
&lt;li&gt;each estimator can use a subset of features&lt;/li&gt;
&lt;li&gt;combine their predictions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Again, the goal is to create a robust ensemble rather than rely on one model.&lt;/p&gt;




&lt;h1&gt;
  
  
  91. Scaling experiment in your notebook
&lt;/h1&gt;

&lt;p&gt;You used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;StandardScaler&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and compared KNN/SVR.&lt;/p&gt;

&lt;p&gt;This is a very useful experiment because it demonstrates:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Preprocessing requirements depend on the algorithm.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tree-based algorithms generally do not need feature scaling in the same way that distance- or margin-based methods do.&lt;/p&gt;

&lt;p&gt;KNN and SVM are much more sensitive to feature scale.&lt;/p&gt;




&lt;h1&gt;
  
  
  92. A subtle issue in your scaling experiment
&lt;/h1&gt;

&lt;p&gt;You created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train_scaled&lt;/span&gt;
&lt;span class="n"&gt;X_test_scaled&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;but your KNN code first trained on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train_log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before later looping over scaled KNN.&lt;/p&gt;

&lt;p&gt;That is useful as experimentation, but for a clean article you should present the comparison more systematically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KNN without scaling
        vs
KNN with scaling
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and explicitly explain why the result changed.&lt;/p&gt;

&lt;p&gt;For SVR, you correctly used the scaled data in the shown section.&lt;/p&gt;




&lt;h1&gt;
  
  
  93. The complete mental model
&lt;/h1&gt;

&lt;p&gt;After these five days, the entire workflow can be remembered as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             REAL-WORLD QUESTION
                     ↓
             Define the target
                     ↓
              Collect data
                     ↓
              Understand data
                     ↓
          ┌──────────┴──────────┐
          ↓                     ↓
      Numerical             Categorical
          ↓                     ↓
      distributions        category counts
          └──────────┬──────────┘
                     ↓
                    EDA
                     ↓
              Clean the data
                     ↓
          Handle missing values
                     ↓
             Encode categories
                     ↓
          Select useful features
                     ↓
              Split train/test
                     ↓
          Scale when appropriate
                     ↓
              Train baseline
                     ↓
          Evaluate on test data
                     ↓
             Try other models
                     ↓
         Tune hyperparameters
                     ↓
       Select based on validation
                     ↓
             Final evaluation
                     ↓
              Make predictions
                     ↓
       Integrate into an application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  94. What each notebook taught you
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Day 1 — Auto MPG
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;load a dataset&lt;/li&gt;
&lt;li&gt;inspect rows and columns&lt;/li&gt;
&lt;li&gt;understand shape and data types&lt;/li&gt;
&lt;li&gt;inspect summary statistics&lt;/li&gt;
&lt;li&gt;identify categorical/text columns&lt;/li&gt;
&lt;li&gt;detect messy numerical values&lt;/li&gt;
&lt;li&gt;visualize distributions&lt;/li&gt;
&lt;li&gt;use countplots&lt;/li&gt;
&lt;li&gt;use histograms&lt;/li&gt;
&lt;li&gt;use pairplots&lt;/li&gt;
&lt;li&gt;understand missing values&lt;/li&gt;
&lt;li&gt;convert &lt;code&gt;"?"&lt;/code&gt; to &lt;code&gt;NaN&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;convert text to numerical data&lt;/li&gt;
&lt;li&gt;use median imputation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Before modeling, understand and clean your data.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 2 — Telco Customer Churn
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identify a classification target&lt;/li&gt;
&lt;li&gt;remove an identifier&lt;/li&gt;
&lt;li&gt;handle &lt;code&gt;TotalCharges&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;visualize class balance&lt;/li&gt;
&lt;li&gt;compare categories against churn&lt;/li&gt;
&lt;li&gt;use boxplots and histograms&lt;/li&gt;
&lt;li&gt;use pairplots&lt;/li&gt;
&lt;li&gt;inspect correlations&lt;/li&gt;
&lt;li&gt;encode categorical data&lt;/li&gt;
&lt;li&gt;split into training/testing data&lt;/li&gt;
&lt;li&gt;train Logistic Regression&lt;/li&gt;
&lt;li&gt;train SVC&lt;/li&gt;
&lt;li&gt;train Decision Tree&lt;/li&gt;
&lt;li&gt;compare train/test performance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A classification model predicts categories, but the quality of the prediction depends heavily on preprocessing and evaluation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 3 — Improving the Telco project
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one-hot encode multi-category variables&lt;/li&gt;
&lt;li&gt;use &lt;code&gt;random_state&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;use &lt;code&gt;stratify&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;scale features&lt;/li&gt;
&lt;li&gt;train Logistic Regression&lt;/li&gt;
&lt;li&gt;build a prediction interface with Gradio&lt;/li&gt;
&lt;li&gt;reproduce preprocessing at inference time&lt;/li&gt;
&lt;li&gt;output a human-readable prediction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;An ML model is only one part of an ML application.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Day 4/5 — Customer Lifetime Value
&lt;/h2&gt;

&lt;p&gt;Main lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;formulate a regression problem&lt;/li&gt;
&lt;li&gt;inspect numerical/categorical relationships&lt;/li&gt;
&lt;li&gt;automate EDA&lt;/li&gt;
&lt;li&gt;encode categories&lt;/li&gt;
&lt;li&gt;split train/test&lt;/li&gt;
&lt;li&gt;compare multiple regressors&lt;/li&gt;
&lt;li&gt;understand R²&lt;/li&gt;
&lt;li&gt;try target transformation&lt;/li&gt;
&lt;li&gt;use log transformation&lt;/li&gt;
&lt;li&gt;reverse the transformation&lt;/li&gt;
&lt;li&gt;experiment with scaling&lt;/li&gt;
&lt;li&gt;tune hyperparameters&lt;/li&gt;
&lt;li&gt;use ensemble models&lt;/li&gt;
&lt;li&gt;generate prediction files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The biggest lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Different models capture different patterns, and model selection should be based on performance on unseen data—not training performance alone.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  95. What I would change before publishing the project as "production-quality"
&lt;/h1&gt;

&lt;p&gt;Your notebooks are excellent learning material because they show experimentation. For a public article, however, it is worth separating:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Learning experiments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recommended ML workflow
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here are the most important improvements.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 1 — Use clear train/test terminology
&lt;/h2&gt;

&lt;p&gt;Prefer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;rather than mixing &lt;code&gt;Y&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Python convention generally uses lowercase &lt;code&gt;y&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 2 — Make experiments reproducible
&lt;/h2&gt;

&lt;p&gt;Use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;consistently when appropriate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 3 — Use stratification for classification
&lt;/h2&gt;

&lt;p&gt;For churn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stratify&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Improvement 4 — Avoid LabelEncoder for nominal multi-class input features
&lt;/h2&gt;

&lt;p&gt;Use one-hot encoding for categories such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PaymentMethod
Contract
InternetService
VehicleClass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Label encoding is much more naturally suited to targets or truly ordinal categories.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 5 — Fit preprocessing only on training data
&lt;/h2&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data
    ↓
fit scaler
    ↓
transform training

Test data
    ↓
transform using existing scaler
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not independently fit preprocessing on the test data.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 6 — Use proper evaluation metrics
&lt;/h2&gt;

&lt;p&gt;For classification, consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accuracy
Precision
Recall
F1
Confusion Matrix
ROC-AUC
PR-AUC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For regression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R²
MAE
MSE
RMSE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which metrics matter depends on the business problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 7 — Don't compare only training scores
&lt;/h2&gt;

&lt;p&gt;Always look at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training performance
+
Validation/test performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A very high training score with much lower test performance can indicate overfitting.&lt;/p&gt;




&lt;h2&gt;
  
  
  Improvement 8 — Keep a validation strategy for model selection
&lt;/h2&gt;

&lt;p&gt;A single test set should ideally be kept for final evaluation.&lt;/p&gt;

&lt;p&gt;During experimentation, use cross-validation or a validation set to choose models/hyperparameters.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data
   ↓
Cross-validation / validation
   ↓
Choose model
   ↓
Final test set
   ↓
Final unbiased-ish estimate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  96. A very important distinction: EDA vs ML
&lt;/h1&gt;

&lt;p&gt;EDA asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What does my data look like?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Machine learning asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can I learn a useful relationship that lets me make predictions on new data?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These overlap, but they are not the same.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;scatterplot&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not build a model.&lt;/p&gt;

&lt;p&gt;It helps you understand the data before deciding how to model it.&lt;/p&gt;




&lt;h1&gt;
  
  
  97. Another important distinction: correlation vs prediction
&lt;/h1&gt;

&lt;p&gt;A strong correlation can be useful.&lt;/p&gt;

&lt;p&gt;But:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;high correlation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;does not automatically mean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;excellent predictive model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prediction depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;multiple features&lt;/li&gt;
&lt;li&gt;data quality&lt;/li&gt;
&lt;li&gt;model type&lt;/li&gt;
&lt;li&gt;noise&lt;/li&gt;
&lt;li&gt;generalization&lt;/li&gt;
&lt;li&gt;preprocessing&lt;/li&gt;
&lt;li&gt;evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So don't judge an ML problem from one graph.&lt;/p&gt;




&lt;h1&gt;
  
  
  98. Another important distinction: model vs algorithm
&lt;/h1&gt;

&lt;p&gt;It is common to hear:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I trained a model."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The algorithm is the learning procedure.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Random Forest = algorithm/model family
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After training:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model_rf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you have a fitted model containing learned information.&lt;/p&gt;

&lt;p&gt;A useful mental model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Algorithm
   +
Training data
   ↓
Fitted model
   ↓
Predictions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  99. What "learning" actually means
&lt;/h1&gt;

&lt;p&gt;This is probably the most important concept to understand.&lt;/p&gt;

&lt;p&gt;When you run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model is not magically understanding the world.&lt;/p&gt;

&lt;p&gt;It is optimizing internal parameters so that its predictions match the training examples according to the algorithm's objective.&lt;/p&gt;

&lt;p&gt;For example, linear regression learns coefficients.&lt;/p&gt;

&lt;p&gt;A tree learns split rules.&lt;/p&gt;

&lt;p&gt;A neural network learns weights.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Training means finding model parameters that make the model perform well according to a defined objective.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  100. What happens during prediction?
&lt;/h1&gt;

&lt;p&gt;Once training is complete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prediction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model does not learn again from the test examples.&lt;/p&gt;

&lt;p&gt;It applies what it learned during training.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training:

X_train + y_train
        ↓
      model.fit()
        ↓
   learned model

Prediction:

X_test
   ↓
learned model
   ↓
prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  101. The biggest beginner misconception to avoid
&lt;/h1&gt;

&lt;p&gt;Do not think:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I used Random Forest, so I did machine learning."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The real ML skill is being able to explain:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What problem are you solving?&lt;/li&gt;
&lt;li&gt;What is the target?&lt;/li&gt;
&lt;li&gt;What features are available?&lt;/li&gt;
&lt;li&gt;What does the data look like?&lt;/li&gt;
&lt;li&gt;What problems did you find?&lt;/li&gt;
&lt;li&gt;How did you clean them?&lt;/li&gt;
&lt;li&gt;How did you encode the data?&lt;/li&gt;
&lt;li&gt;Why did you choose the algorithm?&lt;/li&gt;
&lt;li&gt;How did you evaluate it?&lt;/li&gt;
&lt;li&gt;Does it generalize?&lt;/li&gt;
&lt;li&gt;What would you improve?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you can explain these, you are learning ML—not just memorizing scikit-learn commands.&lt;/p&gt;




&lt;h1&gt;
  
  
  102. A simple explanation of your five-day journey
&lt;/h1&gt;

&lt;p&gt;If someone asks:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What did you actually learn in five days?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A strong answer is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I learned that machine learning is not just about selecting an algorithm. I started by understanding datasets using Pandas, cleaning missing and incorrectly formatted values, and using visualization to understand distributions and relationships. Then I learned the difference between classification and regression, converted categorical data into numerical representations, split data into training and testing sets, trained several models using scikit-learn, and compared their performance on unseen data. Finally, I experimented with scaling, transformations, ensemble models, hyperparameters, and even built a simple Gradio interface around a churn model.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a much stronger story than:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I learned Logistic Regression, SVM and Random Forest."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  103. Suggested article structure for dev.to
&lt;/h1&gt;

&lt;p&gt;For your final public article, I recommend this structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Title

1. Why I started learning ML
2. What I thought ML was vs what I learned
3. AI vs ML
4. Types of ML
5. My five-day learning roadmap
6. Day 1 — Understanding and cleaning data
7. EDA and why graphs matter
8. Day 2 — Classification with customer churn
9. Encoding categorical data
10. Train/test split
11. Logistic Regression, SVM and Decision Trees
12. How I evaluated classification models
13. Day 3 — Turning the model into an application
14. Day 4/5 — Regression and Customer Lifetime Value
15. Linear Regression vs KNN vs SVR vs Trees vs Ensembles
16. Log transformation
17. Scaling
18. Hyperparameters
19. What I got wrong / what I would improve
20. What I learned
21. What I plan to learn next
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This will make the article feel like a &lt;strong&gt;journey&lt;/strong&gt;, rather than a textbook.&lt;/p&gt;




&lt;h1&gt;
  
  
  104. Recommended article narrative
&lt;/h1&gt;

&lt;p&gt;The most interesting story is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"I started with almost no ML knowledge."

        ↓

"I learned that data preparation matters."

        ↓

"I discovered that graphs help me ask questions
before training models."

        ↓

"I learned classification."

        ↓

"I learned regression."

        ↓

"I tried different algorithms and realized
there is no single best model."

        ↓

"I learned that a high training score
doesn't necessarily mean a good model."

        ↓

"I turned one model into a small application."

        ↓

"I now understand the basic ML workflow
and know what I need to learn next."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a much more authentic beginner-to-ML story.&lt;/p&gt;




&lt;h1&gt;
  
  
  105. What to learn next
&lt;/h1&gt;

&lt;p&gt;Based on what your notebooks already cover, I would &lt;strong&gt;not&lt;/strong&gt; recommend immediately jumping into deep learning.&lt;/p&gt;

&lt;p&gt;First strengthen these foundations:&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Statistics
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;mean&lt;/li&gt;
&lt;li&gt;median&lt;/li&gt;
&lt;li&gt;variance&lt;/li&gt;
&lt;li&gt;standard deviation&lt;/li&gt;
&lt;li&gt;distributions&lt;/li&gt;
&lt;li&gt;probability&lt;/li&gt;
&lt;li&gt;correlation&lt;/li&gt;
&lt;li&gt;conditional probability&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. ML evaluation
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;confusion matrix&lt;/li&gt;
&lt;li&gt;precision&lt;/li&gt;
&lt;li&gt;recall&lt;/li&gt;
&lt;li&gt;F1&lt;/li&gt;
&lt;li&gt;ROC-AUC&lt;/li&gt;
&lt;li&gt;MAE&lt;/li&gt;
&lt;li&gt;RMSE&lt;/li&gt;
&lt;li&gt;R²&lt;/li&gt;
&lt;li&gt;cross-validation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Feature engineering
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;handling dates&lt;/li&gt;
&lt;li&gt;categorical features&lt;/li&gt;
&lt;li&gt;transformations&lt;/li&gt;
&lt;li&gt;interactions&lt;/li&gt;
&lt;li&gt;outliers&lt;/li&gt;
&lt;li&gt;missing values&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Model selection
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;cross-validation&lt;/li&gt;
&lt;li&gt;hyperparameter tuning&lt;/li&gt;
&lt;li&gt;GridSearchCV&lt;/li&gt;
&lt;li&gt;RandomizedSearchCV&lt;/li&gt;
&lt;li&gt;pipelines&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. ML fundamentals
&lt;/h2&gt;

&lt;p&gt;Learn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;bias vs variance&lt;/li&gt;
&lt;li&gt;underfitting&lt;/li&gt;
&lt;li&gt;overfitting&lt;/li&gt;
&lt;li&gt;regularization&lt;/li&gt;
&lt;li&gt;data leakage&lt;/li&gt;
&lt;li&gt;feature importance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then move toward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Classical ML
     ↓
Feature engineering
     ↓
Model selection
     ↓
Cross-validation
     ↓
ML pipelines
     ↓
Deployment
     ↓
Deep Learning
     ↓
Generative AI / LLMs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  106. The one-page cheat sheet
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Problem type
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Predict a category → Classification
Predict a number   → Regression
Find hidden groups → Clustering
Learn through reward → Reinforcement Learning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Data
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Rows       → observations
Columns    → variables/features
X          → input features
y          → target
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  First inspection
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sample&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnull&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Visualization
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Countplot   → categories
Histogram   → numerical distribution
Boxplot     → spread/outliers
Scatterplot → two numerical variables
Pairplot    → many numerical relationships
Heatmap     → correlation matrix
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Preparation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Missing values
Data types
Categorical encoding
Feature selection
Scaling when needed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Training
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Prediction
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Evaluation
&lt;/h2&gt;

&lt;p&gt;Classification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Accuracy
Precision
Recall
F1
Confusion Matrix
ROC-AUC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Regression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;R²
MAE
MSE
RMSE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Core warning signs
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score &amp;gt;&amp;gt; Test score
        ↓
Possible overfitting

Test information used during preprocessing
        ↓
Possible data leakage

Identifier used as feature
        ↓
Possible meaningless pattern

Categorical values converted to arbitrary numbers
        ↓
Possible false ordering
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  107. Final takeaway
&lt;/h1&gt;

&lt;p&gt;The biggest lesson from these five days is not a particular Python library or algorithm.&lt;/p&gt;

&lt;p&gt;It is the workflow:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Understand the problem → understand the data → clean the data → explore the data → prepare features → train → evaluate → improve → deploy.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The algorithms are tools inside that workflow.&lt;/p&gt;

&lt;p&gt;You have already touched many of the important building blocks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pandas&lt;/li&gt;
&lt;li&gt;NumPy&lt;/li&gt;
&lt;li&gt;Matplotlib&lt;/li&gt;
&lt;li&gt;Seaborn&lt;/li&gt;
&lt;li&gt;Scikit-learn&lt;/li&gt;
&lt;li&gt;data cleaning&lt;/li&gt;
&lt;li&gt;EDA&lt;/li&gt;
&lt;li&gt;classification&lt;/li&gt;
&lt;li&gt;regression&lt;/li&gt;
&lt;li&gt;encoding&lt;/li&gt;
&lt;li&gt;train/test split&lt;/li&gt;
&lt;li&gt;scaling&lt;/li&gt;
&lt;li&gt;model evaluation&lt;/li&gt;
&lt;li&gt;hyperparameters&lt;/li&gt;
&lt;li&gt;ensemble models&lt;/li&gt;
&lt;li&gt;target transformation&lt;/li&gt;
&lt;li&gt;Gradio&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next step is not to memorize more algorithms.&lt;/p&gt;

&lt;p&gt;The next step is to become comfortable answering:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why am I doing this step?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you can answer that question for every important line in your notebook, you are moving from "following an ML tutorial" to actually understanding machine learning.&lt;/p&gt;




&lt;h1&gt;
  
  
  Appendix A — Notebook-by-notebook step map
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;11_Aug_mpg.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Notebook step&lt;/th&gt;
&lt;th&gt;Why it was done&lt;/th&gt;
&lt;th&gt;ML concept&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Import NumPy/Pandas/Matplotlib/Seaborn&lt;/td&gt;
&lt;td&gt;Load tools&lt;/td&gt;
&lt;td&gt;Python ML stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;read_csv()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Load data&lt;/td&gt;
&lt;td&gt;Data ingestion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nunique()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand uniqueness&lt;/td&gt;
&lt;td&gt;Data exploration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;head()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect examples&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;shape&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand dataset size&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;describe()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Summary statistics&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;info()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Types/non-null values&lt;/td&gt;
&lt;td&gt;Data quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;countplot()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect categorical distribution&lt;/td&gt;
&lt;td&gt;Visualization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hist()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect numerical distributions&lt;/td&gt;
&lt;td&gt;Visualization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;isnull()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Detect actual missing values&lt;/td&gt;
&lt;td&gt;Data cleaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pairplot()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Explore relationships&lt;/td&gt;
&lt;td&gt;EDA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sample()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect random examples&lt;/td&gt;
&lt;td&gt;Data quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identify &lt;code&gt;car name&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Recognize identifier-like field&lt;/td&gt;
&lt;td&gt;Feature selection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replace &lt;code&gt;?&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Convert fake missing values&lt;/td&gt;
&lt;td&gt;Data cleaning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;to_numeric()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Make horsepower numerical&lt;/td&gt;
&lt;td&gt;Data preparation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median imputation&lt;/td&gt;
&lt;td&gt;Fill missing horsepower&lt;/td&gt;
&lt;td&gt;Missing-value handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train/test split&lt;/td&gt;
&lt;td&gt;Separate learning/evaluation data&lt;/td&gt;
&lt;td&gt;Generalization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Logistic Regression&lt;/td&gt;
&lt;td&gt;Start classification experiment&lt;/td&gt;
&lt;td&gt;Supervised learning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Important caveat
&lt;/h3&gt;

&lt;p&gt;The notebook switches from Auto MPG to a &lt;code&gt;loan_prediction.csv&lt;/code&gt; dataset near the end. That appears to be a copied/reused modeling experiment rather than a continuation of the Auto MPG workflow. For a polished article, present these as separate experiments rather than one continuous Auto MPG pipeline.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;code&gt;12th-Aug-Telco-Customer-churn.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Load Telco CSV&lt;/td&gt;
&lt;td&gt;Get customer data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;head()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect records&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;shape&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Understand size&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;describe()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Statistical overview&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drop &lt;code&gt;customerID&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Remove identifier from features&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;info()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Check data types&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sample()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Inspect random rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot of Churn&lt;/td&gt;
&lt;td&gt;Check class distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countplot of gender&lt;/td&gt;
&lt;td&gt;Explore category distribution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Histograms&lt;/td&gt;
&lt;td&gt;Explore numerical distributions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pairplot&lt;/td&gt;
&lt;td&gt;Explore relationships&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Convert TotalCharges&lt;/td&gt;
&lt;td&gt;Fix numeric type&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fill missing TotalCharges&lt;/td&gt;
&lt;td&gt;Handle missing data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LabelEncoder&lt;/td&gt;
&lt;td&gt;Convert categories to numbers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train/test split&lt;/td&gt;
&lt;td&gt;Evaluate on unseen data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Logistic Regression&lt;/td&gt;
&lt;td&gt;Classification model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SVC&lt;/td&gt;
&lt;td&gt;Alternative classifier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision Tree&lt;/td&gt;
&lt;td&gt;Alternative classifier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gradio&lt;/td&gt;
&lt;td&gt;Turn model into interactive app&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Important caveat
&lt;/h3&gt;

&lt;p&gt;The notebook uses LabelEncoder on many categorical input columns. This is useful for learning encoding, but one-hot encoding is generally safer for nominal multi-category features.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;code&gt;12-Aug-2.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;This notebook is a more developed version of the Telco project.&lt;/p&gt;

&lt;p&gt;Important improvements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dropping &lt;code&gt;customerID&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;converting &lt;code&gt;TotalCharges&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;one-hot encoding multi-category variables&lt;/li&gt;
&lt;li&gt;using &lt;code&gt;random_state=42&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;using &lt;code&gt;stratify=Y&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;scaling features&lt;/li&gt;
&lt;li&gt;training Logistic Regression&lt;/li&gt;
&lt;li&gt;building a Gradio prediction interface&lt;/li&gt;
&lt;li&gt;preserving training feature columns&lt;/li&gt;
&lt;li&gt;applying the same preprocessing to new input&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the notebook that most clearly demonstrates the transition from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ML experiment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;small ML application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  &lt;code&gt;13-Aug.ipynb&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Main steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load AutoInsurance
       ↓
Inspect dataset
       ↓
Drop Customer / date columns
       ↓
EDA
       ↓
Identify numerical/categorical columns
       ↓
Generate many visualizations
       ↓
Encode categorical variables
       ↓
Set CLV as target
       ↓
Train/test split
       ↓
Try multiple regressors
       ↓
Try log target
       ↓
Try scaling
       ↓
Tune tree/ensemble hyperparameters
       ↓
Generate predictions
       ↓
Save prediction CSVs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This notebook demonstrates the transition from classification to regression and from a single baseline model to model comparison and tuning.&lt;/p&gt;




&lt;h1&gt;
  
  
  Appendix B — A cleaner ML template to remember
&lt;/h1&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# 1. Load
&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_csv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 2. Understand
&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;head&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# 3. Clean
# handle missing values
# fix data types
# remove inappropriate identifiers
&lt;/span&gt;
&lt;span class="c1"&gt;# 4. Explore
# histograms
# boxplots
# countplots
# scatterplots
# correlation
&lt;/span&gt;
&lt;span class="c1"&gt;# 5. Prepare
&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;target&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# 6. Split
&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_test&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 7. Preprocess
# encoding / scaling
&lt;/span&gt;
&lt;span class="c1"&gt;# 8. Train
&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 9. Predict
&lt;/span&gt;&lt;span class="n"&gt;predictions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# 10. Evaluate
# choose metrics appropriate to the problem
&lt;/span&gt;
&lt;span class="c1"&gt;# 11. Compare / improve
# try another model
# tune hyperparameters
# use cross-validation
&lt;/span&gt;
&lt;span class="c1"&gt;# 12. Final model
# train using the selected approach
&lt;/span&gt;
&lt;span class="c1"&gt;# 13. Predict new data
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  Closing thought for the article
&lt;/h1&gt;

&lt;p&gt;Five days ago, an ML notebook could easily look like a collection of unfamiliar commands.&lt;/p&gt;

&lt;p&gt;Now there is a structure behind those commands.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;head()&lt;/code&gt; is not just a command — it is a way to understand the data.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;isnull()&lt;/code&gt; is not just a command — it is part of data quality.&lt;/p&gt;

&lt;p&gt;A histogram is not just a graph — it tells you about a distribution.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;train_test_split()&lt;/code&gt; is not just boilerplate — it protects your evaluation from simply measuring memorization.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;fit()&lt;/code&gt; is not magic — it is the learning stage.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;predict()&lt;/code&gt; is the model applying what it learned.&lt;/p&gt;

&lt;p&gt;And model comparison is not about finding the fanciest algorithm — it is about finding an approach that generalizes well to data the model has never seen.&lt;/p&gt;

&lt;p&gt;That, more than anything else, is what I took away from my first five days of machine learning.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>ML Foundations: A Complete Data Cleaning &amp; ML Pipeline</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Tue, 25 Aug 2026 13:42:00 +0000</pubDate>
      <link>https://dev.to/yashnigam/ml-foundations-a-complete-data-cleaning-ml-pipeline-2o7o</link>
      <guid>https://dev.to/yashnigam/ml-foundations-a-complete-data-cleaning-ml-pipeline-2o7o</guid>
      <description>&lt;h1&gt;
  
  
  Fundamentals of Machine Learning
&lt;/h1&gt;

&lt;h3&gt;
  
  
  Using Google Colab and Python Libraries to Run ML Models
&lt;/h3&gt;

&lt;h3&gt;
  
  
  Find all the jupiter notebooks and datasets here: &lt;a href="https://github.com/yash-nigam/AI-ML-Foundations" rel="noopener noreferrer"&gt;https://github.com/yash-nigam/AI-ML-Foundations&lt;/a&gt;
&lt;/h3&gt;




&lt;h2&gt;
  
  
  Table of Contents
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;What is Artificial Intelligence?&lt;/li&gt;
&lt;li&gt;
What is Machine Learning?

&lt;ul&gt;
&lt;li&gt;2.1 Types of Machine Learning
&lt;/li&gt;
&lt;li&gt;2.2 Supervised Learning
&lt;/li&gt;
&lt;li&gt;2.3 Unsupervised Learning
&lt;/li&gt;
&lt;li&gt;2.4 Reinforcement Learning
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
How Python Helps Implement Machine Learning

&lt;ul&gt;
&lt;li&gt;3.1 The Libraries and What Each One Handles
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Exploratory Data Analysis: Investigating the Dataset&lt;/li&gt;
&lt;li&gt;
Working with the Loan Prediction Dataset

&lt;ul&gt;
&lt;li&gt;5.1 Making the Dataset Ready for ML Models
&lt;/li&gt;
&lt;li&gt;5.2 What Should Be Cleaned Up
&lt;/li&gt;
&lt;li&gt;5.3 Cleanup Steps

&lt;ul&gt;
&lt;li&gt;5.3.1 Remove Unique/Identifier Columns
&lt;/li&gt;
&lt;li&gt;5.3.2 Handle Missing Values
&lt;/li&gt;
&lt;li&gt;5.3.3 Standardize &amp;amp; Encode Categorical Values
&lt;/li&gt;
&lt;li&gt;5.3.4 Outliers, Duplicates, Imbalance, Scaling, Split
&lt;/li&gt;
&lt;li&gt;5.3.5 Replace Non-Standard Placeholder Values
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Encoding

&lt;ul&gt;
&lt;li&gt;6.1 Why Does ML Need Encoding?
&lt;/li&gt;
&lt;li&gt;6.2 Label Encoding
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Train/Test Split

&lt;ul&gt;
&lt;li&gt;7.1 Why Not Train and Test on the Same Data?
&lt;/li&gt;
&lt;li&gt;7.2 Underfitting
&lt;/li&gt;
&lt;li&gt;7.3 Why &lt;code&gt;random_state&lt;/code&gt; Matters
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Machine Learning Models

&lt;ul&gt;
&lt;li&gt;8.1 Confusion Matrix
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gradio: Turning a Model into an Application&lt;/li&gt;
&lt;li&gt;
What "Learning" Actually Means

&lt;ul&gt;
&lt;li&gt;10.1 What Happens During Prediction?
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The Real ML Skill — Questions to Be Able to Answer&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  1. What is Artificial Intelligence?
&lt;/h2&gt;

&lt;p&gt;Artificial Intelligence (AI) is the broad idea of building systems that can perform tasks that normally require some form of human intelligence, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Recognizing an image&lt;/li&gt;
&lt;li&gt;Generating text / writing an email&lt;/li&gt;
&lt;li&gt;Generating an image&lt;/li&gt;
&lt;li&gt;Recommending a movie&lt;/li&gt;
&lt;li&gt;Detecting fraud&lt;/li&gt;
&lt;li&gt;Predicting the price/value of something&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;A useful mental model for the hierarchy:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiw8x453jzek52q6gcn3p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiw8x453jzek52q6gcn3p.png" alt="AI Hierarchy" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  2. What is Machine Learning?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Machine learning is teaching a computer to find patterns in examples, instead of telling it exact rules to follow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; The old way to predict if a customer will cancel their telecom subscription — you could try writing rules by hand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF contract = month-to-month
AND monthly charges = high
AND tenure = low
THEN churn = yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The machine learning way:&lt;/strong&gt; show the ML algorithm many past customers along with what actually happened to them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Customer&lt;/th&gt;
&lt;th&gt;Contract type&lt;/th&gt;
&lt;th&gt;Monthly charges&lt;/th&gt;
&lt;th&gt;Tenure&lt;/th&gt;
&lt;th&gt;Did they churn?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;Month-to-month&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;Two year&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;Month-to-month&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;Month-to-month&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The algorithm studies this data and figures out the pattern on its own — no one told it "low tenure + high charges = risky."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using it to predict for a new customer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once trained, you can hand it a brand-new customer it has never seen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;New customer's info
        ↓
   Trained model
        ↓
  Predicted: churn or not
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2.1 Types of Machine Learning
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Supervised Learning&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Unsupervised Learning&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Reinforcement Learning&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core idea&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Learn from examples with known answers&lt;/td&gt;
&lt;td&gt;Find hidden patterns with no answers given&lt;/td&gt;
&lt;td&gt;Learn by trial and error, guided by rewards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Data needed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Inputs + labels&lt;/td&gt;
&lt;td&gt;Inputs only, no labels&lt;/td&gt;
&lt;td&gt;No fixed dataset — agent generates its own data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Task types&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Classification, Regression&lt;/td&gt;
&lt;td&gt;Clustering, dimensionality reduction, anomaly detection&lt;/td&gt;
&lt;td&gt;Choosing actions to maximize long-term reward&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Common algorithms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Logistic Regression, Decision Trees, Random Forest&lt;/td&gt;
&lt;td&gt;K-Means, DBSCAN, PCA&lt;/td&gt;
&lt;td&gt;Q-Learning, DQN, Policy Gradient&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;How it's evaluated&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Accuracy, precision, recall, MAE, RMSE&lt;/td&gt;
&lt;td&gt;No ground truth — silhouette score, human judgment&lt;/td&gt;
&lt;td&gt;Total reward earned&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Real-world examples&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Churn prediction, spam filtering, price estimation&lt;/td&gt;
&lt;td&gt;Customer segmentation, topic discovery&lt;/td&gt;
&lt;td&gt;Game-playing agents, robotics, ad bidding&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  2.2 Supervised Learning
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84acxtl1dtmwla9h9x5g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84acxtl1dtmwla9h9x5g.png" alt="Supervised Learning Hierarchy" width="800" height="729"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Supervised learning is the case where &lt;strong&gt;you already know the right answers for your training data.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model makes a guess, checks it against the known answer, sees how wrong it was, and adjusts. Repeat this thousands of times and it gets good at guessing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input (X)  →  Known answer (Y)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;X is everything you know about a customer. Y is what actually happened to them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;X = [tenure=2, contract=month-to-month, charges=95.5]   →   Y = churned
X = [tenure=48, contract=two-year, charges=45.2]        →   Y = stayed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model's job is to learn the &lt;em&gt;relationship&lt;/em&gt; between X and Y well enough that when a new X shows up with no Y attached, it can produce a sensible guess.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The key requirement:&lt;/strong&gt; you need historical data where the outcome is already recorded. No answer key, no supervised learning.&lt;/p&gt;

&lt;h4&gt;
  
  
  The only question that splits supervised learning in two
&lt;/h4&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Is the thing I'm predicting a category, or a number?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's it. That single question decides whether you're doing classification or regression — and it changes your algorithms, your metrics, and how you evaluate success.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Classification&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Regression&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Question it answers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Which category?&lt;/td&gt;
&lt;td&gt;How much?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Type of answer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A fixed label (bucket)&lt;/td&gt;
&lt;td&gt;A number on a scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Examples&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Spam or Not spam&lt;br&gt;Churn or No churn&lt;br&gt;Cat / Dog / Horse&lt;/td&gt;
&lt;td&gt;House price: ₹87,45,000&lt;br&gt;Temperature: 31.4°C&lt;br&gt;Fuel efficiency: 23.7 MPG&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What the output means&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A label, not a quantity — &lt;code&gt;1&lt;/code&gt; isn't "more" than &lt;code&gt;0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A real quantity — 6,000 truly is twice 3,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model output&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A probability → converted to a label&lt;br&gt;e.g. &lt;code&gt;0.91 → "Churn"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A direct number&lt;br&gt;e.g. &lt;code&gt;6,500&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Threshold involved?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes — usually 0.5, adjustable&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Being "wrong"&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Binary — right label or wrong label&lt;/td&gt;
&lt;td&gt;A matter of degree — off by a little or a lot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Common metrics&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Accuracy, Precision, Recall&lt;/td&gt;
&lt;td&gt;MAE, RMSE, R²&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Common algorithms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Logistic Regression, SVC, Decision Tree Classifier&lt;/td&gt;
&lt;td&gt;Linear Regression, KNN Regressor, SVR, Random Forest Regressor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;In your project&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Telco churn (Yes/No)&lt;/td&gt;
&lt;td&gt;Customer Lifetime Value (a number)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Quick test for your own problems&lt;/strong&gt; — when you get a new dataset, look at the target column and ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Does averaging two values in this column produce something meaningful?

  Average of ₹5,000 and ₹7,000 = ₹6,000  ✓ meaningful  → regression
  Average of "spam" and "not spam"       ✗ meaningless → classification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2.3 Unsupervised Learning
&lt;/h3&gt;

&lt;p&gt;In unsupervised learning, we do not have a known target. Instead, the algorithm tries to discover structure in the data.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer data
     ↓
Find groups
     ↓
Group 1: high-value customers
Group 2: price-sensitive customers
Group 3: new customers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Common examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;K-Means clustering&lt;/li&gt;
&lt;li&gt;Hierarchical clustering&lt;/li&gt;
&lt;li&gt;PCA&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2.4 Reinforcement Learning
&lt;/h3&gt;

&lt;p&gt;Reinforcement learning is different. An agent interacts with an environment and receives rewards or penalties.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓ action
Environment
  ↓
Reward / penalty
  ↓
Agent learns
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Game-playing agents&lt;/li&gt;
&lt;li&gt;Robotics&lt;/li&gt;
&lt;li&gt;Certain recommendation/control systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This was not part of the five-day practical work.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  3. How Python Helps Implement Machine Learning
&lt;/h2&gt;

&lt;p&gt;Machine learning is mostly mathematics — matrix operations, optimization, statistics. In principle you could write all of it yourself, but in practice nobody does. Python has become the default language for ML because a small set of mature libraries already implement those pieces, tested and optimized, so you can focus on the problem rather than the arithmetic.&lt;/p&gt;

&lt;p&gt;What makes Python work well here is that these libraries fit together as a &lt;strong&gt;pipeline&lt;/strong&gt;, not as isolated tools. Each one covers one stage of the journey, and they hand data to each other in a common format — a Pandas DataFrame or a NumPy array — so nothing needs translating in between.&lt;/p&gt;

&lt;p&gt;A typical project moves through the stack like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load and clean the data        →  Pandas (with NumPy underneath)
Explore and visualize it       →  Matplotlib + Seaborn
Prepare features and split     →  Scikit-learn
Train and predict              →  Scikit-learn
Evaluate the results           →  Scikit-learn (+ Seaborn to plot them)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that roughly the first half of that pipeline has nothing to do with modeling at all. In real projects, loading, cleaning, and understanding the data usually takes far more time than calling &lt;code&gt;.fit()&lt;/code&gt;. The libraries reflect that reality: Pandas and Seaborn get used constantly, while the actual model training is often two lines.&lt;/p&gt;

&lt;p&gt;The other thing Python gives you is a &lt;strong&gt;consistent interface&lt;/strong&gt;. Once you learn scikit-learn's &lt;code&gt;fit → predict&lt;/code&gt; pattern, it works the same way for logistic regression, random forests, and support vector machines alike. Swapping one algorithm for another is a one-line change, which makes it cheap to try several and compare.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.1 The Libraries and What Each One Handles
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Library&lt;/th&gt;
&lt;th&gt;Import&lt;/th&gt;
&lt;th&gt;Role in the Pipeline&lt;/th&gt;
&lt;th&gt;Key Functions/Methods&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;NumPy&lt;/td&gt;
&lt;td&gt;&lt;code&gt;import numpy as np&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The numerical foundation — fast array math that every other library is built on&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;np.nan&lt;/code&gt;, &lt;code&gt;np.log1p()&lt;/code&gt;, &lt;code&gt;np.expm1()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;np.nan&lt;/code&gt; = missing value. &lt;code&gt;log1p(x)&lt;/code&gt; = log(1+x), useful for skewed targets. &lt;code&gt;expm1(x)&lt;/code&gt; reverses it — used to convert log predictions back to the original CLV scale. You rarely use NumPy directly; it works underneath Pandas and scikit-learn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pandas&lt;/td&gt;
&lt;td&gt;&lt;code&gt;import pandas as pd&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Load and clean the data — a spreadsheet Python can manipulate&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;read_csv()&lt;/code&gt;, column selection/removal, dtype conversion, missing-value detection/fill, X/Y selection, encoding, saving predictions&lt;/td&gt;
&lt;td&gt;Where most of your time actually goes. Example: &lt;code&gt;df = pd.read_csv("Telco_Customer_Churn.csv")&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Matplotlib&lt;/td&gt;
&lt;td&gt;&lt;code&gt;import matplotlib.pyplot as plt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The plotting engine — controls figures, titles, axes, display&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;plt.figure()&lt;/code&gt;, &lt;code&gt;plt.title()&lt;/code&gt;, &lt;code&gt;plt.show()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Rarely used alone; it's the layer Seaborn sits on top of&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seaborn&lt;/td&gt;
&lt;td&gt;&lt;code&gt;import seaborn as sns&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Explore and understand the data through statistical plots&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;countplot&lt;/code&gt;, &lt;code&gt;boxplot&lt;/code&gt;, &lt;code&gt;histplot&lt;/code&gt;, &lt;code&gt;pairplot&lt;/code&gt;, &lt;code&gt;scatterplot&lt;/code&gt;, &lt;code&gt;heatmap&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;One line gives you a plot that would take many in raw Matplotlib. A graph is a tool for asking questions about the data, not decoration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scikit-learn&lt;/td&gt;
&lt;td&gt;&lt;code&gt;from sklearn... import ...&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Everything modeling-related: splitting, preprocessing, scaling, encoding, training, prediction, evaluation&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;train_test_split()&lt;/code&gt;, &lt;code&gt;StandardScaler()&lt;/code&gt;, &lt;code&gt;model.fit(X_train, Y_train)&lt;/code&gt;, &lt;code&gt;model.predict(X_test)&lt;/code&gt;, metrics functions&lt;/td&gt;
&lt;td&gt;The &lt;code&gt;fit → predict&lt;/code&gt; pattern is identical across nearly every model, so trying a different algorithm is a one-line change&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The short version:&lt;/strong&gt; Pandas gets the data into shape, Seaborn helps you understand it, and scikit-learn learns from it — with NumPy doing the arithmetic underneath and Matplotlib drawing the pictures.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Exploratory Data Analysis: Investigating the Dataset
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;What It Does &amp;amp; How to Use It&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;df.head()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Shows the first few rows of the dataset. Run this right after loading data to confirm it loaded correctly and get a quick look at the columns and values.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;df.shape&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Returns the number of rows and columns, e.g. &lt;code&gt;(398, 9)&lt;/code&gt;. Use it to quickly gauge dataset size, which affects model choice and computation time.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;df.info()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lists each column's data type and non-null count. Use it to spot missing values and type problems — e.g. a numeric column showing as &lt;code&gt;object&lt;/code&gt; usually means it contains hidden text like &lt;code&gt;"?"&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;df.describe()&lt;/code&gt; / &lt;code&gt;df.describe(include="all")&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Shows summary statistics (count, mean, min, max, etc.) for numeric columns, or all columns with &lt;code&gt;include="all"&lt;/code&gt;. Use it to check the scale of each variable and spot possible outliers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;df.sample(n)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Shows &lt;code&gt;n&lt;/code&gt; random rows instead of just the first few. Use it after &lt;code&gt;head()&lt;/code&gt; to catch inconsistencies or unusual values that only appear later in the data.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;df.isnull().sum()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Counts missing (&lt;code&gt;NaN&lt;/code&gt;) values per column. Use it to identify which columns need cleaning — note it only catches true &lt;code&gt;NaN&lt;/code&gt;s, not disguised placeholders like &lt;code&gt;"?"&lt;/code&gt;, which need to be checked separately.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  5. Working with the Loan Prediction Dataset
&lt;/h2&gt;

&lt;h3&gt;
  
  
  5.1 Making the Dataset Ready for ML Models
&lt;/h3&gt;

&lt;p&gt;Raw data collected from forms, surveys, or real-world systems is almost never ready to feed straight into a machine learning model.&lt;/p&gt;

&lt;p&gt;ML models cannot by themselves guess around these problems, and these gaps could reduce the accuracy of your results.&lt;/p&gt;

&lt;h3&gt;
  
  
  5.2 What Should Be Cleaned Up
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NaN&lt;/code&gt; (missing values)&lt;/strong&gt; — ML models cannot handle these by themselves; they'll throw errors or silently produce garbage results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unique-identifier columns&lt;/strong&gt; — Unique IDs create noise and aren't helpful for learning genuine patterns in the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text/categorical values&lt;/strong&gt; — Numbers are expected by most ML models, not strings like "Yes"/"No".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Imbalanced or skewed data&lt;/strong&gt; — If 70% of your data is one outcome, a lazy model can get 70% accuracy by always predicting that outcome, without actually learning anything useful.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Outliers and skewed numeric columns&lt;/strong&gt; — A few extreme values can pull the mean far from what's "typical," making mean-based decisions misleading.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5.3 Cleanup Steps
&lt;/h3&gt;

&lt;h4&gt;
  
  
  5.3.1 Remove Unique/Identifier Columns
&lt;/h4&gt;

&lt;p&gt;Columns like IDs (e.g. &lt;code&gt;Loan_ID&lt;/code&gt;) don't carry predictive signal and should be dropped.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Loan_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;axis&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Removes the &lt;code&gt;Loan_ID&lt;/code&gt; column. It's a unique identifier for each row (like a primary key) — it has no predictive value and would only confuse or overfit a model if left in.&lt;/p&gt;

&lt;h4&gt;
  
  
  5.3.2 Handle Missing Values
&lt;/h4&gt;

&lt;p&gt;Either drop rows/columns with too many nulls, or fill them in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;strong&gt;median&lt;/strong&gt; for numeric columns that are skewed (robust to outliers, e.g. income, loan amount).&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;mean&lt;/strong&gt; for numeric columns that are roughly symmetric/normally distributed.&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;mode&lt;/strong&gt; (most frequent value) for categorical columns, since averaging text categories isn't meaningful.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Using Median&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;median_loanamount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LoanAmount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Calculates the median of &lt;code&gt;LoanAmount&lt;/code&gt;. Median is preferred over mean here because loan/income data is typically skewed by a few very high values, and median is more robust to outliers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;median_loanamount_term&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Loan_Amount_Term&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same logic — computes the median loan term to use for filling its missing values later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;median_credit_history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Credit_History&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Computes the median of &lt;code&gt;Credit_History&lt;/code&gt; (a 0/1 flag). Even though it's binary, median is used instead of mean to avoid producing a non-integer/ambiguous fill value.&lt;/p&gt;

&lt;p&gt;Use the calculated values to fill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LoanAmount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LoanAmount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;median_loanamount&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Loan_Amount_Term&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Loan_Amount_Term&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;median_loanamount_term&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Credit_History&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Credit_History&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;median_credit_history&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fills missing values with the calculated medians. This avoids dropping rows (losing data) while keeping the columns usable for modeling, since most ML algorithms can't handle &lt;code&gt;NaN&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using Mode&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Calculates the &lt;strong&gt;mode&lt;/strong&gt; (most frequent value) for gender, married, dependents, and self_employed. Mode is used instead of median/mean because these are categorical columns — you can't average "Male" and "Female".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;mode_gender&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Gender&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;mode_married&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Married&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;mode_dependents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dependents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;mode_self_employed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Self_Employed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace the missing entries with the most common category, which is a reasonable, low-bias guess for a categorical field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Gender&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Gender&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mode_gender&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Married&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Married&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mode_married&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dependents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Dependents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mode_dependents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Self_Employed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Self_Employed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mode_self_employed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  5.3.3 Standardize &amp;amp; Encode Categorical Values
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standardize categorical values&lt;/strong&gt; — fix inconsistent labels (e.g. "male", "Male", "MALE" should all become one consistent value).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encode categorical variables&lt;/strong&gt; — convert text categories into numbers using techniques like Label Encoding or One-Hot Encoding, since most ML models only accept numeric input.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nc"&gt;LabelEncoder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# + loop over cat_cols
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Converts all text/categorical columns (&lt;code&gt;Loan_Status&lt;/code&gt;, &lt;code&gt;Married&lt;/code&gt;, &lt;code&gt;Gender&lt;/code&gt;, etc.) into numeric codes (e.g., "Yes"/"No" → 1/0). This step is &lt;strong&gt;required&lt;/strong&gt; because scikit-learn models only accept numeric input, not strings.&lt;/p&gt;

&lt;h4&gt;
  
  
  5.3.4 Outliers, Duplicates, Imbalance, Scaling, Split
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Check and handle outliers&lt;/strong&gt; — identify extreme values (via boxplots, &lt;code&gt;.describe()&lt;/code&gt;, or z-scores) and decide whether to cap, remove, or transform them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check for and remove duplicate rows&lt;/strong&gt; — duplicate records can bias the model toward those repeated patterns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Address class imbalance&lt;/strong&gt; — if one outcome dominates the target column, apply techniques like upsampling, downsampling, or SMOTE so the model doesn't just learn to predict the majority class.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feature scaling (when needed)&lt;/strong&gt; — normalize or standardize numeric ranges for models sensitive to scale (like SVM, KNN, or Logistic Regression with regularization).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split before further transformation&lt;/strong&gt; — separate train/test sets early so cleaning decisions (like fill values from training data) don't leak information from the test set.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  5.3.5 Replace Non-Standard Placeholder Values
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This converts the placeholder &lt;code&gt;"?"&lt;/code&gt; into a real missing value.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;horsepower&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This then converts the column into a proper numeric type so it can be used in calculations and models.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Encoding
&lt;/h2&gt;

&lt;h3&gt;
  
  
  6.1 Why Does ML Need Encoding?
&lt;/h3&gt;

&lt;p&gt;Many ML algorithms work with numbers. But real-world data contains categories like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Male
Female

Yes
No

Month-to-month
One year
Two year

Fiber optic
DSL
No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model needs these values represented numerically. This is called &lt;strong&gt;encoding&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  6.2 Label Encoding
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.preprocessing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LabelEncoder&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a binary column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can be reasonable for a binary variable. For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Churn:
No  → 0
Yes → 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That makes intuitive sense.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Train/Test Split
&lt;/h2&gt;

&lt;p&gt;Repeatedly used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the most important ML concepts.&lt;/p&gt;

&lt;p&gt;Suppose we have 1,000 examples. We might use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;700 → training
300 → testing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The training set is used to learn the model. The test set is held back to evaluate how the trained model performs on unseen data.&lt;/p&gt;

&lt;p&gt;Think of it like an exam:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training data = practice questions
Test data     = unseen exam questions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  7.1 Why Not Train and Test on the Same Data?
&lt;/h3&gt;

&lt;p&gt;Because the model could simply memorize the training examples. Suppose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 99%
Test score     = 65%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a warning sign. The model learned the training data extremely well but does not generalize well. This is called:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Overfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  7.2 Underfitting
&lt;/h3&gt;

&lt;p&gt;The opposite can happen. If the model is too simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training score = 60%
Test score     = 58%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It may not have learned enough from the data. This is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Underfitting&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A useful mental picture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Underfitting
Model too simple
       ↓
misses important patterns

Good fit
Learns useful patterns
       ↓
works on unseen data

Overfitting
Model learns noise/details
       ↓
great training performance
poor unseen performance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  7.3 Why &lt;code&gt;random_state&lt;/code&gt; Matters
&lt;/h3&gt;

&lt;p&gt;In one Telco notebook:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the improved notebook:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;train_test_split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;test_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;random_state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A random state makes the split reproducible. Without it, the random split can change between runs — meaning your score may change between runs too. With &lt;code&gt;random_state=42&lt;/code&gt;, you can reproduce the same split every time. The number &lt;code&gt;42&lt;/code&gt; is not magical; any fixed integer can serve this purpose.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Machine Learning Models
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model / Concept&lt;/th&gt;
&lt;th&gt;What It Does &amp;amp; How to Use It&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Logistic Regression&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Despite the name, used for classification — estimates the probability of belonging to a class, then applies a threshold to assign a label. A simple, fast, interpretable baseline model — good starting point before trying complex models.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SVC (Support Vector Machine)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Finds a decision boundary that best separates classes, aiming for maximum margin between them. Sensitive to feature scale, so scale your data (e.g. with &lt;code&gt;StandardScaler&lt;/code&gt;) before using it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Decision Tree&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Predicts by asking a sequence of yes/no questions, ending at a leaf with the prediction. Easy to visualize but prone to overfitting — control this with &lt;code&gt;max_depth&lt;/code&gt;, &lt;code&gt;min_samples_split&lt;/code&gt;, &lt;code&gt;min_samples_leaf&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Random Forest&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Combines many decision trees and averages their predictions for a more robust result. An ensemble method — generally more reliable than a single decision tree.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bagging&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Trains multiple models on different random samples of the data, then combines their predictions. Reduces variance and makes predictions more stable; Random Forest is a specialized version of this.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AdaBoost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Builds models sequentially, where each new model focuses on the mistakes of the previous one. A boosting method — improves weak learners step by step rather than averaging independent ones.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KNN (K-Nearest Neighbors)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Predicts by finding the most similar existing examples ("neighbors") and using their values. Relies on distance, so it needs feature scaling to avoid large-scale features dominating the result.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;StandardScaler&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Scales features to comparable ranges before training distance-sensitive models. Always &lt;code&gt;fit_transform&lt;/code&gt; on training data, then only &lt;code&gt;transform&lt;/code&gt; (not fit) on test data, to avoid data leakage.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Linear Regression&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Models the relationship between features and a numeric target as a straight-line equation (&lt;code&gt;Y = b0 + b1X1 + ...&lt;/code&gt;). Used for predicting continuous values, e.g. Customer Lifetime Value.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;R² Score (&lt;code&gt;.score()&lt;/code&gt;)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Measures how well the model explains variation in the target — 1 is perfect, 0 means no better than predicting the average. Report it as "R² of 0.90," not "90% accuracy" — they're not the same thing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Accuracy Score&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The proportion of correct predictions out of all predictions made, used for classification. Easy to understand, but misleading when classes are imbalanced (e.g. 95% "no churn" data).&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  8.1 Confusion Matrix
&lt;/h3&gt;

&lt;p&gt;A confusion matrix organizes classification predictions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Actual
                 No       Yes
Predicted No     TN       FN
Predicted Yes    FP       TP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Meaning:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;True Positive&lt;/strong&gt; — model predicted positive and it was positive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;True Negative&lt;/strong&gt; — model predicted negative and it was negative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;False Positive&lt;/strong&gt; — model predicted positive but it was negative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;False Negative&lt;/strong&gt; — model predicted negative but it was positive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is much more informative than accuracy alone when the cost of errors differs.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. Gradio: Turning a Model into an Application
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;gradio&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;gr&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Building a UI around your model is an important step because it changes the project from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Notebook experiment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User input
    ↓
Preprocessing
    ↓
Model
    ↓
Prediction
    ↓
Human-readable result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gender: Female
Tenure: 12 months
Contract: Month-to-month
Monthly charges: $70
...
          ↓
    Predict Churn
          ↓
High Churn Risk
Probability: ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This demonstrates an important ML engineering concept:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A model is useful only when it can be integrated into a workflow where people or systems can use its predictions.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  10. What "Learning" Actually Means
&lt;/h2&gt;

&lt;p&gt;When you run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_train&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y_train&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model is not magically understanding the world. It is optimizing internal parameters so that its predictions match the training examples according to the algorithm's objective.&lt;/p&gt;

&lt;p&gt;For example, linear regression learns coefficients. A tree learns split rules. A neural network learns weights.&lt;/p&gt;

&lt;p&gt;So:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Training means finding model parameters that make the model perform well according to a defined objective.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  10.1 What Happens During Prediction?
&lt;/h3&gt;

&lt;p&gt;Once training is complete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prediction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;X_test&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the model does not learn again from the test examples. It applies what it learned during training.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Training:

X_train + y_train
        ↓
      model.fit()
        ↓
   learned model

Prediction:

X_test
   ↓
learned model
   ↓
prediction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  11. The Real ML Skill — Questions to Be Able to Answer
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;What problem are you solving?&lt;/li&gt;
&lt;li&gt;What is the target?&lt;/li&gt;
&lt;li&gt;What features are available?&lt;/li&gt;
&lt;li&gt;What does the data look like?&lt;/li&gt;
&lt;li&gt;What problems did you find?&lt;/li&gt;
&lt;li&gt;How did you clean them?&lt;/li&gt;
&lt;li&gt;How did you encode the data?&lt;/li&gt;
&lt;li&gt;Why did you choose the algorithm?&lt;/li&gt;
&lt;li&gt;How did you evaluate it?&lt;/li&gt;
&lt;li&gt;Does it generalize?&lt;/li&gt;
&lt;li&gt;What would you improve?&lt;/li&gt;
&lt;li&gt;Why was a specific algorithm chosen?&lt;/li&gt;
&lt;/ol&gt;

</description>
    </item>
    <item>
      <title>Create a free static website using Github Pages and Jekyll</title>
      <dc:creator>yash-nigam</dc:creator>
      <pubDate>Fri, 03 Mar 2023 17:39:15 +0000</pubDate>
      <link>https://dev.to/yashnigam/create-a-free-static-website-using-github-pages-and-jekyll-41a9</link>
      <guid>https://dev.to/yashnigam/create-a-free-static-website-using-github-pages-and-jekyll-41a9</guid>
      <description>&lt;h2&gt;
  
  
  TOC
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Why use GitHub Pages and Jekyll&lt;/li&gt;
&lt;li&gt;Prerequisites&lt;/li&gt;
&lt;li&gt;1: Dev Machine Setup&lt;/li&gt;
&lt;li&gt;2: Use Chirpy Starter Template&lt;/li&gt;
&lt;li&gt;3: Install dependencies&lt;/li&gt;
&lt;li&gt;4: Modify config yaml and preview changes&lt;/li&gt;
&lt;li&gt;5: Using GitHub Actions&lt;/li&gt;
&lt;li&gt;6: Commit Changes to Repo&lt;/li&gt;
&lt;li&gt;7: View the website&lt;/li&gt;
&lt;li&gt;8: Personalize by Adding Avatar, favicon, sidebar link and create first post&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why use GitHub Pages and Jekyll ? &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://pages.github.com/" rel="noopener noreferrer"&gt;GitHub Pages&lt;/a&gt; allows us to host a static website(html, CSS, JavaScript) on GitHub repository for absolutely free which is ideal for personal websites. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;A repository with the naming convention of &lt;strong&gt;{github-user-name}.github.io&lt;/strong&gt; makes that repository serve a static website with subdomain as {github-user-name} and domain name as github.io&lt;/em&gt; &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://jekyllrb.com/" rel="noopener noreferrer"&gt;Jekyll&lt;/a&gt; (static site generator tool) is integrated with GitHub pages providing professional looking themes and many free templates for your website.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When using Jekyll we do not have to code any HTML, CSS or JavaScript. Instead we only edit the markdown files to achieve the desired results in the templates.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Prerequisites &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Windows/Linux developer machine&lt;/li&gt;
&lt;li&gt;GitHub Account&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1: setup Ruby, Gem and Jekyll on developer machine &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If dev machine is windows, use Windows subsystem for Linux - which installs a ubuntu environment which is available from PowerShell terminal.

&lt;ul&gt;
&lt;li&gt;Open a PowerShell prompt as an Administrator and run:
&lt;code&gt;wsl --install&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;li&gt;Install Ruby and other prerequisites:
&lt;/li&gt;

&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github$ sudo apt-get install ruby-full build-essential zlib1g-dev
[sudo] password for yash:
Reading package lists... Done
Building dependency tree... Done
Reading state information... Done
build-essential is already the newest version (12.9ubuntu3).
zlib1g-dev is already the newest version (1:1.2.11.dfsg-2ubuntu9.2).
The following additional packages will be installed:
  ri
The following NEW packages will be installed:
  ri ruby-full
0 upgraded, 2 newly installed, 0 to remove and 4 not upgraded.
Need to get 6788 B of archives.
After this operation, 38.9 kB of additional disk space will be used.
Do you want to continue? [Y/n] Y
Get:1 http://archive.ubuntu.com/ubuntu jammy/universe amd64 ri all 1:3.0~exp1 [4206 B]
Get:2 http://archive.ubuntu.com/ubuntu jammy/universe amd64 ruby-full all 1:3.0~exp1 [2582 B]
Fetched 6788 B in 1s (4970 B/s)
Selecting previously unselected package ri.
(Reading database ... 47201 files and directories currently installed.)
Preparing to unpack .../ri_1%3a3.0~exp1_all.deb ...
Unpacking ri (1:3.0~exp1) ...
Selecting previously unselected package ruby-full.
Preparing to unpack .../ruby-full_1%3a3.0~exp1_all.deb ...
Unpacking ruby-full (1:3.0~exp1) ...
Setting up ri (1:3.0~exp1) ...
Setting up ruby-full (1:3.0~exp1) ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Add environment variables to your ~/.bashrc file
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;echo '# Install Ruby Gems to ~/gems' &amp;gt;&amp;gt; ~/.bashrc
echo 'export GEM_HOME="$HOME/gems"' &amp;gt;&amp;gt; ~/.bashrc
echo 'export PATH="$HOME/gems/bin:$PATH"' &amp;gt;&amp;gt; ~/.bashrc
source ~/.bashrc
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Install Jekyll and Bundler
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github$ gem install jekyll bundler
Fetching jekyll-4.3.2.gem
Successfully installed jekyll-4.3.2
Parsing documentation for jekyll-4.3.2
Installing ri documentation for jekyll-4.3.2
Done installing documentation for jekyll after 3 seconds
Fetching bundler-2.4.7.gem
Successfully installed bundler-2.4.7
Parsing documentation for bundler-2.4.7
Installing ri documentation for bundler-2.4.7
Done installing documentation for bundler after 0 seconds
2 gems installed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  2: Use chirpy-starter template to create new repo &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Sign in to GitHub.com with your account, and go to &lt;a href="https://github.com/cotes2020/chirpy-starter/" rel="noopener noreferrer"&gt;chirpy-starter&lt;/a&gt;, click the button Use this template &amp;gt; Create a new repository&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffujl101qotj754h2761v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ffujl101qotj754h2761v.png" alt="Use Chirpy Starter Template" width="800" height="290"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In the next window type the repository name to be the same as your GitHub user name appended with .github.io
&lt;code&gt;
&amp;lt;GitHub User Name&amp;gt;.githubpages.io
&lt;/code&gt;
&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fql6tf9k43ov5ri57mz97.png" alt="Create New repository(with same name as GH user) from Template" width="800" height="592"&gt;
This makes the repo. act as a source of static website, and the repo. name becomes the domain name website.&lt;/li&gt;
&lt;li&gt;Once this operation is complete all the contents of &lt;a href="https://github.com/cotes2020/chirpy-starter/" rel="noopener noreferrer"&gt;chirpy-starter&lt;/a&gt; will be present in our new site: &lt;a href="https://github.com/yash-nigam/yash-nigam.github.io" rel="noopener noreferrer"&gt;https://github.com/yash-nigam/yash-nigam.github.io&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3: Installing dependencies of chirpy-starter &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;This site requires some GEM's to be installed before it can be used, the list of can be found in the Gemfile: &lt;a href="https://github.com/cotes2020/chirpy-starter/blob/main/Gemfile" rel="noopener noreferrer"&gt;https://github.com/cotes2020/chirpy-starter/blob/main/Gemfile&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;On your local development machine, clone the new repo
&lt;code&gt;git clone https://github.com/yash-nigam/yash-nigam.github.io.git
&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;go to the repo root directory and run bundle command&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;cd yash-nigam.github.io&lt;br&gt;
&lt;/code&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ bundle
Fetching gem metadata from https://rubygems.org/...........
Resolving dependencies...
Using bundler 2.4.7
Using concurrent-ruby 1.2.2
Using rainbow 3.1.1
Using ffi 1.15.5
Using forwardable-extended 2.6.0
Using mercenary 0.4.0
Using parallel 1.22.1
Using rouge 4.1.0
Using eventmachine 1.2.7
Using unicode-display_width 2.4.2
Using racc 1.6.2
Using colorator 1.1.0
Using i18n 1.12.0
Using ethon 0.16.0
Using yell 2.2.2
Using pathutil 0.16.2
Using liquid 4.0.4
Using public_suffix 5.0.1
Using typhoeus 1.4.0
Using addressable 2.8.1
Using webrick 1.8.1
Using jekyll-paginate 1.1.0
Using http_parser.rb 0.8.0
Using rb-inotify 0.10.1
Using rb-fsevent 0.11.2
Using rexml 3.2.5
Using terminal-table 3.0.2
Using nokogiri 1.14.2 (x86_64-linux)
Using safe_yaml 1.0.5
Using google-protobuf 3.22.0 (x86_64-linux)
Using em-websocket 0.5.3
Using listen 3.8.0
Using sass-embedded 1.58.3 (x86_64-linux-gnu)
Using html-proofer 3.19.4
Using kramdown 2.4.0
Using jekyll-watch 2.2.1
Using jekyll-sass-converter 3.0.0
Using kramdown-parser-gfm 1.1.0
Using jekyll 4.3.2
Using jekyll-archives 2.2.1
Using jekyll-seo-tag 2.8.0
Using jekyll-redirect-from 0.16.0
Using jekyll-sitemap 1.4.0
Fetching jekyll-theme-chirpy 5.5.2
Installing jekyll-theme-chirpy 5.5.2
Bundle complete! 6 Gemfile dependencies, 44 gems now installed.
Use `bundle info [gemname]` to see where a bundled gem is installed.
1 installed gem you directly depend on is looking for funding.
  Run `bundle fund` for details
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4: Modify _config.yml and Preview the changes locally &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Open _config.yml and modify the title and url
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;title: Yash Nigam

url: 'https://yash-nigam.github.io'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Preview the site contents locally:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ bundle exec jekyll s
Configuration file: /home/yash/github/yash-nigam.github.io/_config.yml
            Source: /home/yash/github/yash-nigam.github.io
       Destination: /home/yash/github/yash-nigam.github.io/_site
 Incremental build: disabled. Enable with --incremental
      Generating...
                    done in 1.634 seconds.
 Auto-regeneration: enabled for '/home/yash/github/yash-nigam.github.io'
    Server address: http://127.0.0.1:4000/

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;The local service will be published at &lt;a href="http://127.0.0.1:4000" rel="noopener noreferrer"&gt;http://127.0.0.1:4000&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F82ip5bw1u9dtuwftt9sd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F82ip5bw1u9dtuwftt9sd.png" alt="Local Preview" width="800" height="699"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5: Configure to Deploy the Website using GitHub Actions &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Browse to your repository on GitHub, Select Settings &amp;gt; Pages. Then in the Source section (under Build and deployment), select GitHub Actions from the dropdown menu.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1kpz6iqtvjf8bdvdewgv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1kpz6iqtvjf8bdvdewgv.png" alt="Select GitHub Actions" width="800" height="532"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  6: Commit and push changes to repo&lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Check git status
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ git status
On branch main
Your branch is up to date with 'origin/main'.

Changes not staged for commit:
  (use "git add &amp;lt;file&amp;gt;..." to update what will be committed)
  (use "git restore &amp;lt;file&amp;gt;..." to discard changes in working directory)
        modified:   _config.yml

no changes added to commit (use "git add" and/or "git commit -a")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;add, Commit and push the _config.yaml file to repo
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ git add _config.yml
yash@LearningPC:~/github/yash-nigam.github.io$ git commit -m "updated"
[main c30df68] updated
 1 file changed, 2 insertions(+), 2 deletions(-)
yash@LearningPC:~/github/yash-nigam.github.io$ git push
Username for 'https://github.com': yash-nigam
Password for 'https://yash-nigam@github.com':
Enumerating objects: 5, done.
Counting objects: 100% (5/5), done.
Delta compression using up to 12 threads
Compressing objects: 100% (3/3), done.
Writing objects: 100% (3/3), 334 bytes | 334.00 KiB/s, done.
Total 3 (delta 2), reused 0 (delta 0), pack-reused 0
remote: Resolving deltas: 100% (2/2), completed with 2 local objects.
To https://github.com/yash-nigam/yash-nigam.github.io.git
   aae12d1..c30df68  main -&amp;gt; main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  7: Check the Actions Workflow and view the final website &lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Once the changes are pushed to the repo, a build and deploy workflow should be automatically triggered.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdv6tgd2lo23sia372ppc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdv6tgd2lo23sia372ppc.png" alt="Build and Deploy Workflow" width="800" height="826"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;View the final website
&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F913hc7e0538rpk7izevi.png" alt="Final Website" width="800" height="480"&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  8: Add Avatar, favicon, a new post and an sidebar tab.&lt;a&gt;&lt;/a&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Add Avatar and favicon images
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Create a directory named img under assets in the root of repo
&lt;code&gt;yash-nigam.github.io/assets&lt;/code&gt; and favicons under img.&lt;/li&gt;
&lt;li&gt;Follow the steps to generate and add 
favicons&lt;a href="https://chirpy.cotes.page/posts/customize-the-favicon/" rel="noopener noreferrer"&gt;https://chirpy.cotes.page/posts/customize-the-favicon/&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io/assets$ ls -ltr img/
total 96
drwxr-xr-x 2 yash yash  4096 Mar  3 23:28 favicons
-rw-r--r-- 1 yash yash    90 Mar  4  2023 favicon.ico:Zone.Identifier
-rw-r--r-- 1 yash yash 27019 Mar  4  2023 yn_small.jpg
-rw-r--r-- 1 yash yash 53754 Mar  4  2023 sidebar_bg.jpg
yash@LearningPC:~/github/yash-nigam.github.io$ ls -ltr assets/img/favicons/
total 168
-rw-r--r-- 1 yash yash 15531 Mar  4  2023 android-chrome-192x192.png
-rw-r--r-- 1 yash yash  8930 Mar  4  2023 mstile-150x150.png
-rw-r--r-- 1 yash yash   923 Mar  4  2023 favicon-16x16.png
-rw-r--r-- 1 yash yash   608 Mar  4  2023 safari-pinned-tab.svg
-rw-r--r-- 1 yash yash 15086 Mar  4  2023 favicon.ico
-rw-r--r-- 1 yash yash 66336 Mar  4  2023 android-chrome-512x512.png
-rw-r--r-- 1 yash yash 13736 Mar  4  2023 apple-touch-icon.png
-rw-r--r-- 1 yash yash  1270 Mar  4  2023 favicon-32x32.png

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Also put your personal photo to be used as avatar in the img folder.&lt;/li&gt;
&lt;li&gt;In the _config.yml file make following changes
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;avatar: /assets/img/yn_small.jpg

github:
  username: yash-nigam
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Under the _posts folder create a file as follows:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ ls -ltr _posts/
total 4
-rw-r--r-- 1 yash yash 446 Mar  4 00:14 2023-03-04-FirstPage.md
yash@LearningPC:~/github/yash-nigam.github.io$ cat _posts/2023-03-04-FirstPage.md
---
title: "First Post"
date: 2023-03-03 12:00:00 +0530
categories: [TOP_CATEGORIE, SUB_CATEGORIE]
tags: [TAG]    , TAG names should always be lowercase
---

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Under the _tabs folder create another markdown file of your choice.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ ls -ltr _tabs/
total 20
-rw-r--r-- 1 yash yash  47 Mar  3 15:00 tags.md
-rw-r--r-- 1 yash yash  56 Mar  3 15:00 categories.md
-rw-r--r-- 1 yash yash  55 Mar  3 15:00 archives.md
-rw-r--r-- 1 yash yash 194 Mar  3 15:00 about.md
-rw-r--r-- 1 yash yash 284 Mar  4 01:25 devto.md
yash@LearningPC:~/github/yash-nigam.github.io$ cat _tabs/devto.md
---
layout: page
title: dev.to Blog Posts
icon: fas fa-flask
order: 1
---
-----------
&amp;gt; &amp;lt;a href="https://dev.to/yashnigam/create-a-free-static-website-using-github-pages-and-jekyll-41a9" target="_blank"&amp;gt;create a personal website using Github Pages with the Jekyll Framework&amp;lt;/a&amp;gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Verify the site locally and view the changes in browser
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bundle exec jekyll s

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnbzpm3ooj4oaki0k4vri.png" alt="changes" width="800" height="605"&gt;
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Commit and push the changes
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ git status
On branch main
Your branch is up to date with 'origin/main'.

Changes not staged for commit:
  (use "git add &amp;lt;file&amp;gt;..." to update what will be committed)
  (use "git restore &amp;lt;file&amp;gt;..." to discard changes in working directory)
        modified:   _config.yml

Untracked files:
  (use "git add &amp;lt;file&amp;gt;..." to include in what will be committed)
        _posts/2023-03-04-FirstPage.md
        _tabs/devto.md
        assets/css/
        assets/img/

no changes added to commit (use "git add" and/or "git commit -a")

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ignore any Zone.Identifier files as they are autogenerated&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;yash@LearningPC:~/github/yash-nigam.github.io$ git add --all
warning: CRLF will be replaced by LF in _tabs/devto.md.
The file will have its original line endings in your working directory

yash@LearningPC:~/github/yash-nigam.github.io$ git commit -m "Added changes"
[main 6eded15] Added changes
 24 files changed, 64 insertions(+), 8 deletions(-)
 create mode 100644 _posts/2023-03-04-FirstPage.md
 create mode 100644 _tabs/devto.md
 create mode 100644 assets/css/style.scss
 create mode 100644 assets/img/favicons/android-chrome-192x192.png
 create mode 100644 assets/img/favicons/android-chrome-512x512.png
 create mode 100644 assets/img/favicons/apple-touch-icon.png
 create mode 100644 assets/img/favicons/favicon-16x16.png
 create mode 100644 assets/img/favicons/favicon-32x32.png
 create mode 100644 assets/img/favicons/favicon.ico
 create mode 100644 assets/img/favicons/mstile-150x150.png
 create mode 100644 assets/img/favicons/safari-pinned-tab.svg
 create mode 100644 assets/img/sidebar_bg.jpg
 create mode 100644 assets/img/yn_small.jpg

yash@LearningPC:~/github/yash-nigam.github.io$ git push
remote: Resolving deltas: 100% (3/3), completed with 3 local objects.
To https://github.com/yash-nigam/yash-nigam.github.io.git
   c30df68..6eded15  main -&amp;gt; main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Once changes are pushed, make sure GitHub build and deploy action is successful &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;https://github.com/yash-nigam/yash-nigam.github.io/actions&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Added changes&lt;br&gt;
Build and Deploy #2: Commit 6eded15 pushed by yash-nigam&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Now in the actual deployed site we can see that favicon, avatar, side bar tab and a new blog post is created successfully.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0r1j003fep6fyimi6rtl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0r1j003fep6fyimi6rtl.png" alt="Final Site" width="800" height="689"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  References
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://chirpy.cotes.page/posts/getting-started/" rel="noopener noreferrer"&gt;https://chirpy.cotes.page/posts/getting-started/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jekyllrb.com/docs/installation/" rel="noopener noreferrer"&gt;https://jekyllrb.com/docs/installation/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ubuntu.com/tutorials/install-ubuntu-on-wsl2-on-windows-11-with-gui-support#2-install-wsl" rel="noopener noreferrer"&gt;Install Windows Subsystem for Linux&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://chirpy.cotes.page/posts/getting-started/#option-1-using-the-chirpy-starter" rel="noopener noreferrer"&gt;Chirpy Starter&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/pages/getting-started-with-github-pages/configuring-a-publishing-source-for-your-github-pages-site#publishing-with-a-custom-github-actions-workflow" rel="noopener noreferrer"&gt;Publishing with a custom GitHub Actions workflow
&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://chirpy.cotes.page/" rel="noopener noreferrer"&gt;https://chirpy.cotes.page/&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  chirpy implementations
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://jamstackthemes.dev/demo/theme/jekyll-theme-chirpy/" rel="noopener noreferrer"&gt;https://jamstackthemes.dev/demo/theme/jekyll-theme-chirpy/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://bpostance.github.io/" rel="noopener noreferrer"&gt;https://bpostance.github.io/&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>fastapi</category>
      <category>django</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
