<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Tools in Data Science</title><link>/2026-05/</link><description>Recent content on Tools in Data Science</description><generator>Hugo</generator><language>en-US</language><atom:link href="/2026-05/index.xml" rel="self" type="application/rss+xml"/><item><title/><link>/2026-05/actor-network-visualization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/actor-network-visualization/</guid><description>&lt;h2 id="actor-network-visualization"&gt;Actor Network Visualization&lt;a class="anchor" href="#actor-network-visualization"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Find the shortest path between Govinda &amp;amp; Angelina Jolie using IMDb data using Python: &lt;a href="https://pypi.org/project/networkx/"&gt;networkx&lt;/a&gt; or &lt;a href="https://pypi.org/project/scikit-network"&gt;scikit-network&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/lcwMsPxPIjc"&gt;&lt;img src="https://i.ytimg.com/vi_webp/lcwMsPxPIjc/sddefault.webp" alt="Jolie No. 1" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/sanand0/jolie-no-1/blob/master/jolie-no-1.ipynb"&gt;Notebook: How this video was created&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/sanand0/jolie-no-1/blob/master/imdb-actor-pairing.ipynb"&gt;The data used to visualize the network&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/sanand0/jolie-no-1/blob/master/shortest-path.ipynb"&gt;The shortest path between actors&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.imdb.com/non-commercial-datasets/"&gt;IMDb data&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/sanand0/jolie-no-1"&gt;Codebase&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to analyze and visualize the social network of actors using IMDb data, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Acquisition and Preparation:&lt;/strong&gt; Learn how to download and work with IMDb&amp;rsquo;s non-commercial datasets. This includes using Python to read and process large TSV files.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Filtering and Wrangling:&lt;/strong&gt; Understand how to clean and filter large datasets to focus on relevant information, such as specific actor categories (actors/actresses) and movie titles, to make the data more manageable for analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Actor Pairing and Network Creation:&lt;/strong&gt; Discover how to identify significant collaborations between actors by setting thresholds for the minimum number of films and co-appearances. You will learn to create a network of actors based on these pairings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Data Inconsistencies:&lt;/strong&gt; Learn techniques to manage real-world data issues, such as variations in actor names, to ensure the accuracy of your network.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Network Analysis with Python:&lt;/strong&gt; Get introduced to powerful Python libraries for network analysis like &lt;code&gt;networkx&lt;/code&gt; and &lt;code&gt;scikit-network&lt;/code&gt; to build, manipulate, and study the structure of complex networks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Finding the Shortest Path:&lt;/strong&gt; Learn how to apply graph algorithms to find the shortest path between two actors in the network, demonstrating the &amp;ldquo;degrees of separation&amp;rdquo; concept.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Jupyter Notebook for Data Storytelling:&lt;/strong&gt; See how to use Jupyter Notebooks to document and present a data analysis project, combining code, text, and visualizations to tell a compelling story.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exporting Results:&lt;/strong&gt; Learn how to export your findings, such as the actor pairings, into formats like Excel for further use or presentation.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/ai-coding-cli/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-cli/</guid><description>&lt;h1 id="ai-coding-in-the-cli"&gt;AI Coding in the CLI&lt;a class="anchor" href="#ai-coding-in-the-cli"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Command-line AI coding tools bring the power of large language models directly to your terminal, enabling scripted automation, pipeline integration, and efficient developer workflows. These tools excel at batch processing, git integration, and system-level automation that GUI tools can&amp;rsquo;t match.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude Code CLI&lt;/strong&gt;: Terminal-first coding agent for interactive development and review&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Codex CLI (OpenAI)&lt;/strong&gt;: Local coding agent with sandbox, approvals, CI-friendly exec&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GitHub Copilot CLI&lt;/strong&gt;: GitHub-native terminal agent with policy-aware approvals and MCP extensions&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemini CLI&lt;/strong&gt;: Large-context, multimodal agent with non-interactive mode&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Simon Willison&amp;rsquo;s &lt;code&gt;llm&lt;/code&gt;&lt;/strong&gt;: Shell-native AI with UNIX pipeline integration and extensive tooling&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="claude-code-cli"&gt;Claude Code CLI&lt;a class="anchor" href="#claude-code-cli"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://claude.ai/code"&gt;Claude Code&lt;/a&gt; is Anthropic’s terminal-first coding agent (npm: &lt;code&gt;@anthropic-ai/claude-code&lt;/code&gt;). It runs locally in your project, understands your codebase, and assists with routine tasks, explanations, and git workflows.&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-context/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-context/</guid><description>&lt;h1 id="ai-coding-context-engineering"&gt;AI Coding Context Engineering&lt;a class="anchor" href="#ai-coding-context-engineering"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Context engineering is the systematic approach to providing AI coding assistants with the right information at the right level of detail. Unlike general prompt engineering, code context engineering focuses specifically on code-related workflows, project structure, and development processes.&lt;/p&gt;
&lt;p&gt;Effective context engineering transforms AI from a simple code generator into an intelligent development partner that understands your project&amp;rsquo;s architecture, conventions, and goals.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-ide/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-ide/</guid><description>&lt;h1 id="ai-coding-in-ides"&gt;AI Coding in IDEs&lt;a class="anchor" href="#ai-coding-in-ides"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;AI-powered coding assistants have changed how we write code within traditional IDEs and editors. These tools provide intelligent code completion, chat interfaces, and automated refactoring capabilities that dramatically increase developer productivity.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Copy-paste workflows&lt;/strong&gt;: Fast iteration cycles using web-based AI assistants for quick fixes and code generation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GitHub Copilot&lt;/strong&gt;: The industry-standard AI pair programmer for VS Code and other editors&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cursor&lt;/strong&gt;: An AI-native editor with advanced context understanding and agent capabilities&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Windsurf&lt;/strong&gt;: The first agentic IDE with deep codebase awareness and collaborative AI features&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="copy-paste-workflows"&gt;Copy-paste workflows&lt;a class="anchor" href="#copy-paste-workflows"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The simplest and often most effective AI coding workflow involves copying code and errors to web-based AI assistants, then pasting the improved solutions back into your editor. This approach works with any editor and provides access to the latest AI models.&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-online/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-online/</guid><description>&lt;h1 id="ai-coding-online"&gt;AI Coding Online&lt;a class="anchor" href="#ai-coding-online"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Modern AI coding tools have transformed how we build software, enabling rapid prototyping and deployment through natural language interfaces. This module covers the mindset shift toward &amp;ldquo;vibe-coding&amp;rdquo; and the essential tools for AI-assisted development.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Inline AI coding sandboxes&lt;/strong&gt;: Use Claude Artifacts, Gemini Canvas, and ChatGPT Canvas for iterative development&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hosted agent workbenches&lt;/strong&gt;: Leverage Replit Agent, Bolt.new, and Lovable for end-to-end app creation&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="artifacts-inline-ai-coding-sandboxes"&gt;Artifacts: Inline AI coding sandboxes&lt;a class="anchor" href="#artifacts-inline-ai-coding-sandboxes"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;AI artifacts provide side-by-side code and chat interfaces that revolutionize how we iterate on applications. These tools show visible diffs, run code in real-time, and allow for conversational editing.&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-strategies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-strategies/</guid><description>&lt;h1 id="ai-coding-strategies"&gt;AI Coding Strategies&lt;a class="anchor" href="#ai-coding-strategies"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Strategic approaches to AI-assisted software development have evolved rapidly in 2025. Unlike simple code completion, these methodologies focus on sustainable, team-oriented workflows that maximize AI productivity while maintaining code quality and developer control.&lt;/p&gt;
&lt;p&gt;Modern AI coding strategies combine human expertise with AI capabilities through structured patterns, parallel processing, and iterative feedback loops. The most effective approaches treat AI as a collaborative partner rather than a replacement, establishing clear boundaries and verification processes.&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-tests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-tests/</guid><description>&lt;h1 id="ai-coding-tests"&gt;AI Coding Tests&lt;a class="anchor" href="#ai-coding-tests"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Testing AI-generated code is about establishing fast, trustworthy feedback loops before the code lands in main. The goal is not perfect certainty—it is a reliable signal that the patch does what it should, fails when it must, and keeps shipping velocity high.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scorecard-driven evaluation loops&lt;/strong&gt;: Track AI agent success with automated suites and selectors&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Premortem test-first flows&lt;/strong&gt;: Lock in failing tests before asking an agent to code&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Python-first fuzzing and mutation&lt;/strong&gt;: Catch almost-right code with Hypothesis and mutmut&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent-powered smoke tests&lt;/strong&gt;: Let Playwright bots guard critical user journeys&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="build-evaluation-loops"&gt;Build evaluation loops&lt;a class="anchor" href="#build-evaluation-loops"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Benchmarks stop vibe-coded fixes from quietly regressing. Start with a thin scorecard that runs on every candidate patch, then expand.&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding-tools/</guid><description>&lt;h1 id="ai-coding-tools"&gt;AI Coding Tools&lt;a class="anchor" href="#ai-coding-tools"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;AI coding agents perform best when you stage a predictable toolbox around them. The goal is not to spray every binary into the sandbox, but to focus on fast, auditable utilities that agents can compose for data prep, testing, review, and automation. Keep the happy path linear: expose tools, document how to call them, and wire safe defaults before inviting an agent into the repo.&lt;/p&gt;
&lt;p&gt;In this module, you&amp;rsquo;ll learn:&lt;/p&gt;</description></item><item><title/><link>/2026-05/ai-coding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ai-coding/</guid><description>&lt;h1 id="ai-coding"&gt;AI Coding&lt;a class="anchor" href="#ai-coding"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;AI-assisted development pairs human judgment with generative models across every workspace. This module shows how to set up productive collaborations, feed models the right project context, and ship code you trust.&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll explore:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Creative prompting&lt;/strong&gt; that shapes tone and flow through vibe coding sessions and scenario planning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environment-specific tooling&lt;/strong&gt; for working in browser editors, IDE extensions, and CLI companions without losing momentum.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context engineering practices&lt;/strong&gt; that stage repositories, snippets, and docs so models stay grounded in your codebase.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality safeguards&lt;/strong&gt; that weave in automated tests, structured reviews, and targeted tool calls to keep AI-written code production ready.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;By the end you will know when to lean on AI, how to steer it with precise context, and the guardrails that turn generated snippets into maintainable features.&lt;/p&gt;</description></item><item><title/><link>/2026-05/archive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/archive/</guid><description>&lt;h1 id="archived-content"&gt;Archived content&lt;a class="anchor" href="#archived-content"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="videos"&gt;Videos&lt;a class="anchor" href="#videos"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id="iitm-bs-degree---diploma-level-orientation"&gt;IITM BS Degree - Diploma Level Orientation&lt;a class="anchor" href="#iitm-bs-degree---diploma-level-orientation"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://youtu.be/Dj7X0bQRJSs"&gt;&lt;img src="https://img.youtube.com/vi_webp/Dj7X0bQRJSs/sddefault.webp" alt="IITM BS Degree - Diploma Level Orientation, May 2022" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="tools-in-data-science-orientation"&gt;Tools in Data Science Orientation&lt;a class="anchor" href="#tools-in-data-science-orientation"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://youtu.be/_c_aFQ0ObLo?t=4186"&gt;&lt;img src="https://img.youtube.com/vi_webp/_c_aFQ0ObLo/sddefault.webp" alt="TDS Orientation, May 2022" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="tools-in-data-science-live-sessions"&gt;Tools in Data Science Live Sessions&lt;a class="anchor" href="#tools-in-data-science-live-sessions"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/@se-lr5ff"&gt;&lt;img src="https://img.youtube.com/vi_webp/MQgOy5RNNz0/sddefault.webp" alt="TDS YouTube channel for live sessions" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="tools-in-data-science-playlist"&gt;Tools in Data Science Playlist&lt;a class="anchor" href="#tools-in-data-science-playlist"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/playlist?list=PLZ2ps__7DhBZEOxUBCkv61WHOu8m7ObpE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/3OeMOb7gByE/sddefault.webp" alt="Tools in Data Science Course Playlist , but only 17 videos" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtube.com/playlist?list=PLZ2ps__7DhBZJ2q_hd8ZbDRgOJlB0CZLw&amp;amp;feature=shared"&gt;Tools in Data Science Course Playlist, 66 videos&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="optional-parse--clean-pdf-files-with-tabula"&gt;Optional: Parse &amp;amp; clean PDF files with Tabula&lt;a class="anchor" href="#optional-parse--clean-pdf-files-with-tabula"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/IEusn9HB1sc"&gt;&lt;img src="https://i.ytimg.com/vi_webp/IEusn9HB1sc/sddefault.webp" alt="Parse &amp;amp; clean PDF files with Tabula" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/base64-encoding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/base64-encoding/</guid><description>&lt;h1 id="base-64-encoding"&gt;Base 64 Encoding&lt;a class="anchor" href="#base-64-encoding"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Base64 is a method to convert binary data into ASCII text. It&amp;rsquo;s essential when you need to transmit binary data through text-only channels or embed binary content in text formats.&lt;/p&gt;
&lt;p&gt;Watch this quick explanation of how Base64 works (3 min):&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/8qkxeZmKmOY"&gt;&lt;img src="https://i.ytimg.com/vi_webp/8qkxeZmKmOY/sddefault.webp" alt="What is Base64? (3 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s how it works:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It takes 3 bytes (24 bits) and converts them into 4 ASCII characters&lt;/li&gt;
&lt;li&gt;&amp;hellip; using 64 characters: A-Z, a-z, 0-9, + and / (padding with &lt;code&gt;=&lt;/code&gt; to make the length a multiple of 4)&lt;/li&gt;
&lt;li&gt;There&amp;rsquo;s a URL-safe variant of Base64 that replaces + and / with - and _ to avoid issues in URLs&lt;/li&gt;
&lt;li&gt;Base64 adds ~33% overhead (since every 3 bytes becomes 4 characters)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Common Python operations with Base64:&lt;/p&gt;</description></item><item><title/><link>/2026-05/bash/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bash/</guid><description>&lt;h2 id="terminal-bash"&gt;Terminal: Bash&lt;a class="anchor" href="#terminal-bash"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;UNIX shells are the de facto standard in the data science world and &lt;a href="https://www.gnu.org/software/bash/"&gt;Bash&lt;/a&gt; is the most popular.
This is available by default on Mac and Linux.&lt;/p&gt;
&lt;p&gt;On Windows, install &lt;a href="https://git-scm.com/downloads"&gt;Git Bash&lt;/a&gt; or &lt;a href="https://learn.microsoft.com/en-us/windows/wsl/install"&gt;WSL&lt;/a&gt; to get a UNIX shell.&lt;/p&gt;
&lt;p&gt;Watch this video to install WSL (12 min).&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/X-DHaQLrBi8"&gt;&lt;img src="https://i.ytimg.com/vi_webp/X-DHaQLrBi8/sddefault.webp" alt="How to Install Ubuntu on Windows 10 (WSL) (12 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Watch this video to understand the basics of Bash and UNIX shell commands (75 min).&lt;/p&gt;</description></item><item><title/><link>/2026-05/bbc-weather-api-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bbc-weather-api-with-python/</guid><description>&lt;h2 id="bbc-weather-location-id-with-python"&gt;BBC Weather location ID with Python&lt;a class="anchor" href="#bbc-weather-location-id-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/IafLrvnamAw"&gt;&lt;img src="https://i.ytimg.com/vi_webp/IafLrvnamAw/sddefault.webp" alt="BBC Weather location API with Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to get the location ID of any city from the BBC Weather API &amp;ndash; as a precursor to scraping weather data &amp;ndash; covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Understanding API Calls&lt;/strong&gt;: Learn how backend API calls work when searching for a city on the BBC weather website.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inspecting Web Interactions&lt;/strong&gt;: Use the browser&amp;rsquo;s inspect element feature to track API calls and understand the network activity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracting Location IDs&lt;/strong&gt;: Identify and extract the location ID from the API response using Python.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Using Python Libraries&lt;/strong&gt;: Import and use requests, json, and urlencode libraries to make API calls and process responses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Constructing API URLs&lt;/strong&gt;: Create structured API URLs dynamically with constant prefixes and query parameters using urlencode.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Building Functions&lt;/strong&gt;: Develop a Python function that accepts a city name, constructs the API call, and returns the location ID.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To open the browser Developer Tools on Chrome, Edge, or Firefox, you can:&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/bridge-course-syllabus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/bridge-course-syllabus/</guid><description>&lt;h1 id="tds-bridge-bootcamp--combined-checklistsyllabus"&gt;TDS Bridge Bootcamp — Combined Checklist/Syllabus&lt;a class="anchor" href="#tds-bridge-bootcamp--combined-checklistsyllabus"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;hr&gt;
&lt;h1 id="day-1-post-session-checklist"&gt;Day 1 Post-Session Checklist&lt;a class="anchor" href="#day-1-post-session-checklist"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="setup-day"&gt;Setup Day&lt;a class="anchor" href="#setup-day"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can open a Linux terminal and recognize the shell prompt (&lt;code&gt;user@machine:~$&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can run &lt;code&gt;python --version&lt;/code&gt;, &lt;code&gt;git --version&lt;/code&gt;, and &lt;code&gt;uv --version&lt;/code&gt; without errors&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can open VS Code connected to WSL/Linux using &lt;code&gt;code .&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know what &lt;code&gt;~&lt;/code&gt; means and can navigate to my home directory using &lt;code&gt;cd ~&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can create folders and files using &lt;code&gt;mkdir&lt;/code&gt; and &lt;code&gt;touch&lt;/code&gt;, and list them with &lt;code&gt;ls&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know the difference between &lt;code&gt;.xyz&lt;/code&gt;, &lt;code&gt;./xyz&lt;/code&gt;, and &lt;code&gt;../xyz&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know that in Linux, everything is a file&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can navigate to the C and D drives within WSL&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I understand the difference between &lt;code&gt;&amp;gt;&lt;/code&gt; (overwrite) and &lt;code&gt;&amp;gt;&amp;gt;&lt;/code&gt; (append)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I have a GitHub account and have created the &lt;code&gt;tds-bootcamp&lt;/code&gt; repository&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1 id="day-2-post-session-checklist"&gt;Day 2 Post-Session Checklist&lt;a class="anchor" href="#day-2-post-session-checklist"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="linux--shell-essentials"&gt;Linux &amp;amp; Shell Essentials&lt;a class="anchor" href="#linux--shell-essentials"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I understand what &lt;code&gt;PATH&lt;/code&gt; is and why commands like &lt;code&gt;python&lt;/code&gt; work without full paths&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can navigate the filesystem without clicking — using &lt;code&gt;cd&lt;/code&gt;, &lt;code&gt;ls&lt;/code&gt;, and &lt;code&gt;pwd&lt;/code&gt; only&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can read, search, and inspect files using &lt;code&gt;cat&lt;/code&gt;, &lt;code&gt;head&lt;/code&gt;, &lt;code&gt;tail&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;, and &lt;code&gt;wc&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can edit a file using &lt;code&gt;nano&lt;/code&gt; (open, edit, save, exit)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I understand pipes (&lt;code&gt;|&lt;/code&gt;) and redirection (&lt;code&gt;&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;gt;&amp;gt;&lt;/code&gt;, &lt;code&gt;2&amp;gt;&lt;/code&gt;) and can chain commands&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can set an environment variable in &lt;code&gt;.bashrc&lt;/code&gt; and apply it with &lt;code&gt;source ~/.bashrc&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know the difference between &lt;code&gt;export VAR=value&lt;/code&gt; (available to child processes) and just &lt;code&gt;VAR=value&lt;/code&gt; (shell-local)&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1 id="day-3-post-session-checklist"&gt;Day 3 Post-Session Checklist&lt;a class="anchor" href="#day-3-post-session-checklist"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="vs-code--python-projects-with-uv"&gt;VS Code + Python Projects with uv&lt;a class="anchor" href="#vs-code--python-projects-with-uv"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can run a Python script using &lt;code&gt;uv run script.py&lt;/code&gt; without setting up a local virtual environment&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know where the temporary virtual environment is created when running the &lt;code&gt;uv add --script script.py pandas&lt;/code&gt; command followed by &lt;code&gt;uv run script.py&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can create a new Python project using &lt;code&gt;uv init&lt;/code&gt; and understand what &lt;code&gt;pyproject.toml&lt;/code&gt; is used for&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can add a dependency (e.g., &lt;code&gt;requests&lt;/code&gt;) using &lt;code&gt;uv add&lt;/code&gt; and see it reflected in the lockfile&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can create a traditional virtual environment using &lt;code&gt;uv venv&lt;/code&gt; and know when to use it&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I understand why installing packages globally with &lt;code&gt;pip install&lt;/code&gt; is a bad habit&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can open a project in VS Code, select the correct Python interpreter, and run code from the integrated terminal&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know the difference between a &lt;code&gt;.py&lt;/code&gt; script and a &lt;code&gt;.ipynb&lt;/code&gt; notebook, and when to use each&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1 id="day-4-post-session-checklist"&gt;Day 4 Post-Session Checklist&lt;a class="anchor" href="#day-4-post-session-checklist"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="http-apis--chrome-devtools"&gt;HTTP, APIs &amp;amp; Chrome DevTools&lt;a class="anchor" href="#http-apis--chrome-devtools"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know the 5 core HTTP methods (GET, POST, PUT, PATCH, DELETE) and what each is used for&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can read a status code and know what went wrong — e.g., 401 vs. 403 vs. 404 vs. 500&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can open Chrome DevTools Network tab, find a request, and inspect its headers, payload, and response&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can copy a browser request as a cURL command and run it in the terminal&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can change the &lt;code&gt;User-Agent&lt;/code&gt; in the browser and see the change in the Network tab&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can use &lt;code&gt;curl&lt;/code&gt; to make a GET request with query parameters and a POST request with a JSON body&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I have a running FastAPI app with at least two endpoints (&lt;code&gt;GET /health&lt;/code&gt; and &lt;code&gt;POST /echo&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can test my API using the Swagger UI at &lt;code&gt;/docs&lt;/code&gt; and via &lt;code&gt;curl&lt;/code&gt; from the terminal&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1 id="day-5-post-session-checklist"&gt;Day 5 Post-Session Checklist&lt;a class="anchor" href="#day-5-post-session-checklist"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="git--github-workflow"&gt;Git &amp;amp; GitHub Workflow&lt;a class="anchor" href="#git--github-workflow"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I have set up the basic Git configuration using &lt;code&gt;git config --global user.name &amp;quot;Your Name&amp;quot;&lt;/code&gt;, &lt;code&gt;git config --global user.email &amp;quot;your.email@example.com&amp;quot;&lt;/code&gt;, and set the default branch as main using &lt;code&gt;git config --global init.defaultBranch main&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know GitHub allows only one user account per person, so I have merged my accounts (IITM and personal) into a single unified account&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I understand the three states of a file in Git: working tree → staging → committed&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can run the daily workflow commands &lt;code&gt;git status&lt;/code&gt;, &lt;code&gt;git diff&lt;/code&gt;, and &lt;code&gt;git log&lt;/code&gt; and know what each shows&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know how &lt;code&gt;.gitignore&lt;/code&gt; works and how to use it to ignore files and folders that should not be pushed to the remote repository (e.g., &lt;code&gt;venv&lt;/code&gt;, &lt;code&gt;__pycache__&lt;/code&gt;, &lt;code&gt;.env&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can make a commit: &lt;code&gt;git add&lt;/code&gt; → &lt;code&gt;git commit -m &amp;quot;message&amp;quot;&lt;/code&gt; → &lt;code&gt;git push&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I know what &lt;code&gt;origin&lt;/code&gt; and &lt;code&gt;main&lt;/code&gt; are and can explain them in one sentence each&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can set up SSH key authentication and push to GitHub without entering a password&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can create an annotated tag (&lt;code&gt;git tag -a v0.1.0&lt;/code&gt;) and push it to GitHub&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; I can write a meaningful commit message (not &amp;ldquo;fixed stuff&amp;rdquo; or &amp;ldquo;final.py&amp;rdquo;)&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-1-file-structure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-1-file-structure/</guid><description>&lt;h1 id="day-1--file-structure-windows-vs-linux"&gt;Day 1 — File Structure: Windows vs Linux&lt;a class="anchor" href="#day-1--file-structure-windows-vs-linux"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Understand how files are organized on Linux vs Windows so you never get confused about &amp;ldquo;where my file went&amp;rdquo; or &amp;ldquo;why can&amp;rsquo;t Python find my data.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-big-picture"&gt;The Big Picture&lt;a class="anchor" href="#the-big-picture"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;On &lt;strong&gt;Windows&lt;/strong&gt;, your files live under drive letters:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;C:\Users\alice\Documents\project\data.csv
D:\backup\photos\&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;On &lt;strong&gt;Linux&lt;/strong&gt;, everything starts from a single root &lt;code&gt;/&lt;/code&gt;:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;/home/alice/projects/data.csv
/tmp/scratch.txt&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;There are &lt;strong&gt;no drive letters&lt;/strong&gt; on Linux. Every file on the entire system hangs off &lt;code&gt;/&lt;/code&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-1-installation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-1-installation/</guid><description>&lt;h1 id="day-1--installation-guide"&gt;Day 1 — Installation Guide&lt;a class="anchor" href="#day-1--installation-guide"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Get a fully working development environment — terminal, editor, Git, Python, and &lt;code&gt;uv&lt;/code&gt; — so you can start coding from Day 2 onwards with zero friction.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-you-need-installed-by-the-end"&gt;What You Need Installed by the End&lt;a class="anchor" href="#what-you-need-installed-by-the-end"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Tool&lt;/th&gt;
					&lt;th&gt;Purpose&lt;/th&gt;
					&lt;th&gt;Verify with&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Ubuntu Linux&lt;/strong&gt; (native or WSL2)&lt;/td&gt;
					&lt;td&gt;Your operating system / shell environment&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uname -a&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;VS Code&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Code editor&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;code --version&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Git&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Version control&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;git --version&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Python 3.11+&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Programming language&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;python3 --version&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;uv&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Python package/project manager&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv --version&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="option-a--native-ubuntu-linux"&gt;Option A — Native Ubuntu Linux&lt;a class="anchor" href="#option-a--native-ubuntu-linux"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;If you are already running Ubuntu (22.04 or later), you just need to install the tools.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-1-notes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-1-notes/</guid><description>&lt;h1 id="day-1--setup-day-complete-notes"&gt;Day 1 — Setup Day: Complete Notes&lt;a class="anchor" href="#day-1--setup-day-complete-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h3 id="tds-bridge-bootcamp"&gt;TDS Bridge Bootcamp&lt;a class="anchor" href="#tds-bridge-bootcamp"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal of Day 1:&lt;/strong&gt; Before you write a single line of data science code, you need a working environment you &lt;em&gt;understand&lt;/em&gt;. This day is not just &amp;ldquo;install stuff&amp;rdquo; — it is about building a mental model of where your files live, how your computer runs programs, and why things fail when they do.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="table-of-contents"&gt;Table of Contents&lt;a class="anchor" href="#table-of-contents"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#1-linux-vs-windows--the-mental-model"&gt;Linux vs Windows — The Mental Model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#2-what-is-wsl2-and-how-does-it-work"&gt;What is WSL2 and How Does it Work&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#3-filesystem-anatomy--where-everything-lives"&gt;Filesystem Anatomy — Where Everything Lives&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#4-the-terminal-and-the-shell"&gt;The Terminal and the Shell&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#5-path-and-environment-variables"&gt;PATH and Environment Variables&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#6-installing-your-environment"&gt;Installing Your Environment&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#7-verifying-your-setup"&gt;Verifying Your Setup&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-1-notes/#day-1-final-checklist"&gt;Day 1 Final Checklist&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="1-linux-vs-windows--the-mental-model"&gt;1. Linux vs Windows — The Mental Model&lt;a class="anchor" href="#1-linux-vs-windows--the-mental-model"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id="11-why-this-matters"&gt;1.1 Why This Matters&lt;a class="anchor" href="#11-why-this-matters"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Most Python tutorials assume you just &amp;ldquo;open a terminal and type.&amp;rdquo; But if you are on Windows, there are &lt;strong&gt;two completely different worlds&lt;/strong&gt; on your machine: the Windows world and the Linux world (via WSL2). Confusing them is the #1 source of beginner pain.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-1-path-reading/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-1-path-reading/</guid><description>&lt;h1 id="day-1--path-reading"&gt;Day 1 — Path Reading&lt;a class="anchor" href="#day-1--path-reading"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Master absolute paths, relative paths, and special shortcuts (&lt;code&gt;~&lt;/code&gt;, &lt;code&gt;.&lt;/code&gt;, &lt;code&gt;..&lt;/code&gt;) so you can always tell the system exactly where a file is — or figure out where you are.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-a-path"&gt;What Is a Path?&lt;a class="anchor" href="#what-is-a-path"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A &lt;strong&gt;path&lt;/strong&gt; is the address of a file or folder on your system. Just like a street address tells you where a building is, a path tells the computer where to find a file.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-1-quiz-exercises/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-1-quiz-exercises/</guid><description>&lt;h1 id="day-1-quiz--exercises"&gt;Day 1 Quiz &amp;amp; Exercises&lt;a class="anchor" href="#day-1-quiz--exercises"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="tds-bridge-bootcamp--setup-day"&gt;TDS Bridge Bootcamp — Setup Day&lt;a class="anchor" href="#tds-bridge-bootcamp--setup-day"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Instructions:&lt;/strong&gt; Attempt the MCQs first (no Googling). Then do the exercises in your terminal. The goal is to &lt;em&gt;think&lt;/em&gt; before you type.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="part-a--30-multiple-choice-questions"&gt;Part A — 30 Multiple Choice Questions&lt;a class="anchor" href="#part-a--30-multiple-choice-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Q1.&lt;/strong&gt; You open a terminal and see this prompt: &lt;code&gt;alice@laptop:~$&lt;/code&gt;. What does the &lt;code&gt;~&lt;/code&gt; mean?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; A) The terminal is connected to the internet&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; B) You are in the root directory &lt;code&gt;/&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; C) You are in your home directory&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; D) You are running as administrator&lt;/li&gt;
&lt;/ul&gt;
&lt;details&gt;
&lt;summary&gt;Answer&lt;/summary&gt;
&lt;p&gt;&lt;strong&gt;C&lt;/strong&gt; — &lt;code&gt;~&lt;/code&gt; is the shorthand for your home directory.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-2-basic-script-writing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-2-basic-script-writing/</guid><description>&lt;h1 id="day-2--basic-script-writing"&gt;Day 2 — Basic Script Writing&lt;a class="anchor" href="#day-2--basic-script-writing"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Write your first shell scripts — from simple one-liners to multi-step automation scripts with variables, conditionals, and loops.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-a-shell-script"&gt;What Is a Shell Script?&lt;a class="anchor" href="#what-is-a-shell-script"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A shell script is a text file containing a series of commands that the shell (bash) executes in order. Instead of typing 10 commands one by one, you put them in a file and run it once.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Instead of typing these every day:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd ~/projects/tds
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git pull
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python3 main.py
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;Done!&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Put them in a script and run: ./daily.sh&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;hr&gt;
&lt;h2 id="your-first-script"&gt;Your First Script&lt;a class="anchor" href="#your-first-script"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id="step-1-create-the-file"&gt;Step 1: Create the file&lt;a class="anchor" href="#step-1-create-the-file"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;nano ~/hello.sh&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="step-2-write-the-script"&gt;Step 2: Write the script&lt;a class="anchor" href="#step-2-write-the-script"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;#!/bin/bash
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# My first shell script&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;Hello, World!&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;Today is &lt;/span&gt;&lt;span style="color:#66d9ef"&gt;$(&lt;/span&gt;date&lt;span style="color:#66d9ef"&gt;)&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;You are: &lt;/span&gt;&lt;span style="color:#66d9ef"&gt;$(&lt;/span&gt;whoami&lt;span style="color:#66d9ef"&gt;)&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;You are in: &lt;/span&gt;&lt;span style="color:#66d9ef"&gt;$(&lt;/span&gt;pwd&lt;span style="color:#66d9ef"&gt;)&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="step-3-make-it-executable"&gt;Step 3: Make it executable&lt;a class="anchor" href="#step-3-make-it-executable"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;chmod +x ~/hello.sh&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="step-4-run-it"&gt;Step 4: Run it&lt;a class="anchor" href="#step-4-run-it"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;~/hello.sh
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Hello, World!&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Today is Mon Jan 15 10:30:00 IST 2026&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# You are: alice&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# You are in: /home/alice&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;hr&gt;
&lt;h2 id="the-shebang-line--binbash"&gt;The Shebang Line — &lt;code&gt;#!/bin/bash&lt;/code&gt;&lt;a class="anchor" href="#the-shebang-line--binbash"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The first line of every script should be the &lt;strong&gt;shebang&lt;/strong&gt;:&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-2-notes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-2-notes/</guid><description>&lt;h1 id="day-2--linux--shell-essentials-the-survival-toolkit"&gt;Day 2 — Linux &amp;amp; Shell Essentials: The Survival Toolkit&lt;a class="anchor" href="#day-2--linux--shell-essentials-the-survival-toolkit"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h3 id="tds-bridge-bootcamp"&gt;TDS Bridge Bootcamp&lt;a class="anchor" href="#tds-bridge-bootcamp"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal of Day 2:&lt;/strong&gt; Move from &amp;ldquo;I can type a command if someone tells me what to type&amp;rdquo; → &amp;ldquo;I can navigate, manage files, edit text, and write simple shell scripts independently.&amp;rdquo; These are the skills you will use every single day as a data practitioner.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="table-of-contents"&gt;Table of Contents&lt;a class="anchor" href="#table-of-contents"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#1-navigating-the-filesystem"&gt;Navigating the Filesystem&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#2-working-with-files-and-directories"&gt;Working with Files and Directories&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#3-reading-file-contents"&gt;Reading File Contents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#4-pipes-and-redirection"&gt;Pipes and Redirection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#5-text-editors"&gt;Text Editors&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#6-shell-variables-and-simple-scripts"&gt;Shell Variables and Simple Scripts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#7-getting-help"&gt;Getting Help&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/bridge-course/day-2-notes/#8-lab-2--workspace-creator-and-documenter"&gt;Lab 2 — Workspace Creator and Documenter&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="1-navigating-the-filesystem"&gt;1. Navigating the Filesystem&lt;a class="anchor" href="#1-navigating-the-filesystem"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id="11-pwd--where-am-i"&gt;1.1 &lt;code&gt;pwd&lt;/code&gt; — Where Am I?&lt;a class="anchor" href="#11-pwd--where-am-i"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;pwd
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# /home/alice/projects/tds&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;pwd&lt;/code&gt; = &lt;strong&gt;P&lt;/strong&gt;rint &lt;strong&gt;W&lt;/strong&gt;orking &lt;strong&gt;D&lt;/strong&gt;irectory. Always run this if you are lost.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-2-quiz-exercises/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-2-quiz-exercises/</guid><description>&lt;h1 id="day-2-quiz--exercises"&gt;Day 2 Quiz &amp;amp; Exercises&lt;a class="anchor" href="#day-2-quiz--exercises"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="tds-bridge-bootcamp--linux--shell-essentials"&gt;TDS Bridge Bootcamp — Linux &amp;amp; Shell Essentials&lt;a class="anchor" href="#tds-bridge-bootcamp--linux--shell-essentials"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Instructions:&lt;/strong&gt; Attempt the MCQs without help first. Then do the terminal exercises. The exercises build on each other — do them in order.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="part-a--20-multiple-choice-questions"&gt;Part A — 20 Multiple Choice Questions&lt;a class="anchor" href="#part-a--20-multiple-choice-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Q1.&lt;/strong&gt; You type &lt;code&gt;python&lt;/code&gt; in the terminal and get &amp;ldquo;command not found&amp;rdquo;. Your friend on the same machine has no problem. What is the most likely cause?&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-2-terminal-navigation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-2-terminal-navigation/</guid><description>&lt;h1 id="day-2--terminal-navigation"&gt;Day 2 — Terminal Navigation&lt;a class="anchor" href="#day-2--terminal-navigation"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Move around the Linux filesystem with confidence using &lt;code&gt;pwd&lt;/code&gt;, &lt;code&gt;ls&lt;/code&gt;, &lt;code&gt;cd&lt;/code&gt;, and &lt;code&gt;tree&lt;/code&gt; — so you never feel lost in the terminal.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="where-am-i--pwd"&gt;Where Am I? — &lt;code&gt;pwd&lt;/code&gt;&lt;a class="anchor" href="#where-am-i--pwd"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The first command to learn. Run it whenever you are lost:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;pwd
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# /home/alice/projects/tds&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;pwd&lt;/code&gt; = &lt;strong&gt;P&lt;/strong&gt;rint &lt;strong&gt;W&lt;/strong&gt;orking &lt;strong&gt;D&lt;/strong&gt;irectory. It tells you the absolute path of where your terminal is right now.&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Every command you run operates relative to this location. If you run &lt;code&gt;ls&lt;/code&gt;, it lists what&amp;rsquo;s in this directory. If you run &lt;code&gt;python3 script.py&lt;/code&gt;, it looks for &lt;code&gt;script.py&lt;/code&gt; here.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-2-touch-mkdir-rm-cp-mv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-2-touch-mkdir-rm-cp-mv/</guid><description>&lt;h1 id="day-2--file-operations-touch-mkdir-rm-cp-mv"&gt;Day 2 — File Operations: touch, mkdir, rm, cp, mv&lt;a class="anchor" href="#day-2--file-operations-touch-mkdir-rm-cp-mv"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Master the essential file management commands so you can create, copy, move, rename, and delete files and directories entirely from the terminal.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="creating-files--touch"&gt;Creating Files — &lt;code&gt;touch&lt;/code&gt;&lt;a class="anchor" href="#creating-files--touch"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;touch&lt;/code&gt; creates an empty file (or updates the timestamp of an existing file).&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;touch notes.txt &lt;span style="color:#75715e"&gt;# create one empty file&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;touch file1.txt file2.txt file3.txt &lt;span style="color:#75715e"&gt;# create multiple files&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;touch ~/projects/README.md &lt;span style="color:#75715e"&gt;# create in a specific location&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If the file already exists, &lt;code&gt;touch&lt;/code&gt; updates its &amp;ldquo;last modified&amp;rdquo; time without changing content.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-3-quiz-exercises/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-3-quiz-exercises/</guid><description>&lt;h1 id="day-3-quiz--exercises"&gt;Day 3 Quiz &amp;amp; Exercises&lt;a class="anchor" href="#day-3-quiz--exercises"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="tds-bridge-bootcamp--vs-code--python-projects-with-uv"&gt;TDS Bridge Bootcamp — VS Code + Python Projects with uv&lt;a class="anchor" href="#tds-bridge-bootcamp--vs-code--python-projects-with-uv"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Instructions:&lt;/strong&gt; Attempt MCQs first, then do the exercises in order in your terminal.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="part-a--20-multiple-choice-questions"&gt;Part A — 20 Multiple Choice Questions&lt;a class="anchor" href="#part-a--20-multiple-choice-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Q1.&lt;/strong&gt; You run &lt;code&gt;pip install pandas&lt;/code&gt; without a virtual environment. A week later you start a new project that needs a different version of pandas. What is the likely problem?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; A) pip won&amp;rsquo;t allow installing pandas twice&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; B) Both projects share the same global pandas — changing the version for one breaks the other&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; C) pip automatically creates separate environments per project&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; D) You&amp;rsquo;ll need to uninstall Python and reinstall it&lt;/li&gt;
&lt;/ul&gt;
&lt;details&gt;
&lt;summary&gt;Answer&lt;/summary&gt;
&lt;p&gt;&lt;strong&gt;B&lt;/strong&gt; — Both projects share the same global pandas — changing the version for one breaks the other&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-3-uv-workflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-3-uv-workflow/</guid><description>&lt;h1 id="day-3--uv-workflow"&gt;Day 3 — UV Workflow&lt;a class="anchor" href="#day-3--uv-workflow"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Learn the day-to-day &lt;code&gt;uv&lt;/code&gt; commands to create projects, manage dependencies, run scripts, and keep your Python environment reproducible.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-uv"&gt;What Is UV?&lt;a class="anchor" href="#what-is-uv"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;uv&lt;/strong&gt; is an extremely fast Python package and project manager built by &lt;a href="https://astral.sh/"&gt;Astral&lt;/a&gt; (the same team behind &lt;code&gt;ruff&lt;/code&gt;). It replaces several tools:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Traditional tool&lt;/th&gt;
					&lt;th&gt;What it does&lt;/th&gt;
					&lt;th&gt;uv equivalent&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;pip&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Install packages&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv pip install&lt;/code&gt; / &lt;code&gt;uv add&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;venv&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Create virtual environments&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv venv&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;pyenv&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Manage Python versions&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv python install&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;pipx&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Run CLI tools&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uvx&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;pip-tools&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Lock dependencies&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv lock&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Why uv?&lt;/strong&gt; It&amp;rsquo;s 10-100x faster than pip. A full dependency install that takes pip 30 seconds takes uv under 1 second.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-3-virtual-environments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-3-virtual-environments/</guid><description>&lt;h1 id="day-3--virtual-environments"&gt;Day 3 — Virtual Environments&lt;a class="anchor" href="#day-3--virtual-environments"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Understand why virtual environments exist, how they work, and how to create and use them — so your projects never break each other&amp;rsquo;s dependencies.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-problem-why-virtual-environments"&gt;The Problem: Why Virtual Environments?&lt;a class="anchor" href="#the-problem-why-virtual-environments"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Imagine you have two Python projects:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Project A — needs pandas==1.5.0
Project B — needs pandas==2.1.0&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;If you install both globally (system-wide), only one version of &lt;code&gt;pandas&lt;/code&gt; can exist at a time. Installing one breaks the other.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-3-vscode-setup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-3-vscode-setup/</guid><description>&lt;h1 id="day-3--vs-code-setup"&gt;Day 3 — VS Code Setup&lt;a class="anchor" href="#day-3--vs-code-setup"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Set up VS Code as your primary code editor with the right extensions, terminal integration, and settings so you can write, run, and debug Python efficiently.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-vs-code"&gt;What Is VS Code?&lt;a class="anchor" href="#what-is-vs-code"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Visual Studio Code (VS Code)&lt;/strong&gt; is a free, lightweight code editor by Microsoft. It&amp;rsquo;s the most popular editor among developers because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It works on Windows, macOS, and Linux&lt;/li&gt;
&lt;li&gt;It has a built-in terminal&lt;/li&gt;
&lt;li&gt;It has thousands of extensions for every language&lt;/li&gt;
&lt;li&gt;It connects to WSL2, remote servers, and containers&lt;/li&gt;
&lt;li&gt;It has great Git integration built-in&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;VS Code is &lt;strong&gt;not&lt;/strong&gt; the same as Visual Studio (which is a full IDE). VS Code is lighter and more flexible.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-4-api-testing-curl/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-4-api-testing-curl/</guid><description>&lt;h1 id="day-4--api-testing-with-curl"&gt;Day 4 — API Testing with curl&lt;a class="anchor" href="#day-4--api-testing-with-curl"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Use &lt;code&gt;curl&lt;/code&gt; to send HTTP requests from the terminal — GET, POST, headers, JSON bodies — so you can test and debug APIs without leaving the command line.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-curl"&gt;What Is curl?&lt;a class="anchor" href="#what-is-curl"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;curl&lt;/code&gt; (Client URL) is a command-line tool for making HTTP requests. It comes pre-installed on most Linux systems.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl --version
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# curl 8.x.x ...&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Think of &lt;code&gt;curl&lt;/code&gt; as a browser for your terminal — but instead of rendering a web page, it shows you the raw HTTP response.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-4-chrome-devtools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-4-chrome-devtools/</guid><description>&lt;h1 id="day-4--chrome-devtools"&gt;Day 4 — Chrome DevTools&lt;a class="anchor" href="#day-4--chrome-devtools"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Use Chrome DevTools to inspect network requests, debug API calls, and understand what happens behind the scenes when you load a web page or call an API.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-are-chrome-devtools"&gt;What Are Chrome DevTools?&lt;a class="anchor" href="#what-are-chrome-devtools"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Chrome DevTools is a set of developer tools built into Google Chrome. It lets you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Inspect HTML/CSS elements&lt;/li&gt;
&lt;li&gt;Monitor network requests (API calls, file downloads)&lt;/li&gt;
&lt;li&gt;Debug JavaScript&lt;/li&gt;
&lt;li&gt;Analyze performance&lt;/li&gt;
&lt;li&gt;View cookies and storage&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For this course, the &lt;strong&gt;Network tab&lt;/strong&gt; is the most important tool.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-4-http-fundamentals/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-4-http-fundamentals/</guid><description>&lt;h1 id="day-4--http-fundamentals"&gt;Day 4 — HTTP Fundamentals&lt;a class="anchor" href="#day-4--http-fundamentals"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Understand how the web works at the protocol level — HTTP methods, status codes, headers, and request/response structure — so you can debug APIs and build web services.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-http"&gt;What Is HTTP?&lt;a class="anchor" href="#what-is-http"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;HTTP&lt;/strong&gt; (HyperText Transfer Protocol) is the language that browsers and servers speak. Every time you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Open a web page&lt;/li&gt;
&lt;li&gt;Submit a form&lt;/li&gt;
&lt;li&gt;Fetch data from an API&lt;/li&gt;
&lt;li&gt;Upload a file&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;hellip;an HTTP &lt;strong&gt;request&lt;/strong&gt; goes out and an HTTP &lt;strong&gt;response&lt;/strong&gt; comes back.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-4-quiz-exercises/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-4-quiz-exercises/</guid><description>&lt;h1 id="day-4-quiz--exercises"&gt;Day 4 Quiz &amp;amp; Exercises&lt;a class="anchor" href="#day-4-quiz--exercises"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="tds-bridge-bootcamp--http-apis--chrome-devtools"&gt;TDS Bridge Bootcamp — HTTP, APIs &amp;amp; Chrome DevTools&lt;a class="anchor" href="#tds-bridge-bootcamp--http-apis--chrome-devtools"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Instructions:&lt;/strong&gt; Attempt MCQs first. Then run the exercises — you&amp;rsquo;ll need your terminal, a browser with DevTools, and your FastAPI app running.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="part-a--20-multiple-choice-questions"&gt;Part A — 20 Multiple Choice Questions&lt;a class="anchor" href="#part-a--20-multiple-choice-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Q1.&lt;/strong&gt; You click &amp;ldquo;Submit&amp;rdquo; on a login form. The username and password are sent to the server. Which HTTP method is almost certainly being used?&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-5-git-basic-flow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-5-git-basic-flow/</guid><description>&lt;h1 id="day-5--git-basic-flow"&gt;Day 5 — Git Basic Flow&lt;a class="anchor" href="#day-5--git-basic-flow"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Understand the core Git mental model — working tree, staging area, and commits — so you know exactly what happens at each step and never feel lost.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-git"&gt;What Is Git?&lt;a class="anchor" href="#what-is-git"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Git&lt;/strong&gt; is a version control system. It tracks changes to your files over time, letting you:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;See what changed, when, and by whom&lt;/li&gt;
&lt;li&gt;Undo mistakes by going back to any previous version&lt;/li&gt;
&lt;li&gt;Work on features in parallel without breaking things&lt;/li&gt;
&lt;li&gt;Collaborate with others without overwriting each other&amp;rsquo;s work&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Every professional developer uses Git daily. It is non-negotiable.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-5-github-ssh-login/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-5-github-ssh-login/</guid><description>&lt;h1 id="day-5--github-ssh-login"&gt;Day 5 — GitHub SSH Login&lt;a class="anchor" href="#day-5--github-ssh-login"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Set up SSH key authentication for GitHub so you can push and pull without entering passwords, and understand how public-key cryptography works in practice.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="why-ssh"&gt;Why SSH?&lt;a class="anchor" href="#why-ssh"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When you push code to GitHub, GitHub needs to verify you are who you say you are. There are two main ways:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Method&lt;/th&gt;
					&lt;th&gt;How it works&lt;/th&gt;
					&lt;th&gt;Pros&lt;/th&gt;
					&lt;th&gt;Cons&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;HTTPS + Token&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Password/token sent with each request&lt;/td&gt;
					&lt;td&gt;Easy setup&lt;/td&gt;
					&lt;td&gt;Must manage tokens&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;SSH Keys&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Cryptographic key pair&lt;/td&gt;
					&lt;td&gt;No password needed, more secure&lt;/td&gt;
					&lt;td&gt;One-time setup is longer&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;We recommend SSH&lt;/strong&gt; — once set up, it just works. No tokens to manage or expire.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-5-init-add-commit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-5-init-add-commit/</guid><description>&lt;h1 id="day-5--git-init-add-commit"&gt;Day 5 — Git: init, add, commit&lt;a class="anchor" href="#day-5--git-init-add-commit"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Goal:&lt;/strong&gt; Master the hands-on commands for creating repositories, staging changes, and making commits — the commands you&amp;rsquo;ll use every single day.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="git-init--create-a-new-repository"&gt;&lt;code&gt;git init&lt;/code&gt; — Create a New Repository&lt;a class="anchor" href="#git-init--create-a-new-repository"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Create a project directory and initialize Git:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mkdir my-project
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd my-project
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git init
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Initialized empty Git repository in /home/you/my-project/.git/&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="what-git-init-creates"&gt;What &lt;code&gt;git init&lt;/code&gt; creates&lt;a class="anchor" href="#what-git-init-creates"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ls -la .git/
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# HEAD ← points to current branch&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# config ← repository-level settings&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# hooks/ ← scripts that run on events (commit, push, etc.)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# objects/ ← where Git stores file contents and commits&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# refs/ ← branch and tag pointers&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;blockquote class='book-hint '&gt;
&lt;p&gt;You never need to touch &lt;code&gt;.git/&lt;/code&gt; directly. Git commands manage it for you.&lt;/p&gt;</description></item><item><title/><link>/2026-05/bridge-course/day-5-quiz-exercises/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/bridge-course/day-5-quiz-exercises/</guid><description>&lt;h1 id="day-5-quiz--exercises"&gt;Day 5 Quiz &amp;amp; Exercises&lt;a class="anchor" href="#day-5-quiz--exercises"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="tds-bridge-bootcamp--git--github-workflow"&gt;TDS Bridge Bootcamp — Git &amp;amp; GitHub Workflow&lt;a class="anchor" href="#tds-bridge-bootcamp--git--github-workflow"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Instructions:&lt;/strong&gt; Attempt MCQs first. Then do the exercises in order — they build toward your final submission.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="part-a--20-multiple-choice-questions"&gt;Part A — 20 Multiple Choice Questions&lt;a class="anchor" href="#part-a--20-multiple-choice-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Q1.&lt;/strong&gt; You edit a file &lt;code&gt;app.py&lt;/code&gt;. You run &lt;code&gt;git status&lt;/code&gt; and see it listed under &amp;ldquo;Changes not staged for commit&amp;rdquo;. What does this mean?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; A) The file has been committed to Git&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; B) The file has been staged but not committed&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; C) The file was modified but not yet staged&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; D) The file is new and Git doesn&amp;rsquo;t know about it&lt;/li&gt;
&lt;/ul&gt;
&lt;details&gt;
&lt;summary&gt;Answer&lt;/summary&gt;
&lt;p&gt;&lt;strong&gt;C&lt;/strong&gt; — The file was modified but not yet staged&lt;/p&gt;</description></item><item><title/><link>/2026-05/cleaning-data-with-openrefine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/cleaning-data-with-openrefine/</guid><description>&lt;h2 id="cleaning-data-with-openrefine"&gt;Cleaning Data with OpenRefine&lt;a class="anchor" href="#cleaning-data-with-openrefine"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/zxEtfHseE84"&gt;&lt;img src="https://i.ytimg.com/vi_webp/zxEtfHseE84/sddefault.webp" alt="Cleaning data with OpenRefine" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This session covers the use of OpenRefine for data cleaning, focusing on resolving entity discrepancies:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Upload and Project Creation&lt;/strong&gt;: Import data into OpenRefine and create a new project for analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Faceting Data&lt;/strong&gt;: Use text facets to group similar entries and identify frequency of address crumbs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clustering Methodology&lt;/strong&gt;: Apply clustering algorithms to merge similar entries with minor differences, such as punctuation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Manual and Automated Clustering&lt;/strong&gt;: Learn to merge clusters manually or in one go, trusting the system&amp;rsquo;s clustering accuracy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Entity Resolution&lt;/strong&gt;: Clean and save the data by resolving multiple versions of the same entity using Open Refine.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/colab/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/colab/</guid><description>&lt;h2 id="notebooks-google-colab"&gt;Notebooks: Google Colab&lt;a class="anchor" href="#notebooks-google-colab"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://colab.research.google.com/"&gt;Google Colab&lt;/a&gt; is a free, cloud-based Jupyter notebook environment that&amp;rsquo;s become indispensable for data scientists and ML practitioners. It&amp;rsquo;s particularly valuable because it provides free access to GPUs and TPUs, and for easy sharing of code and execution results.&lt;/p&gt;
&lt;p&gt;While Colab is excellent for prototyping and learning, its free tier has limitations - notebooks time out after 12 hours, and GPU access can be inconsistent.&lt;/p&gt;
&lt;p&gt;Learn how to mount Google Drive for persistent storage, manage dependencies with &lt;code&gt;!pip install&lt;/code&gt; commands, as these are common pain points when getting started.&lt;/p&gt;</description></item><item><title/><link>/2026-05/convert-html-to-markdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/convert-html-to-markdown/</guid><description>&lt;h2 id="converting-html-to-markdown"&gt;Converting HTML to Markdown&lt;a class="anchor" href="#converting-html-to-markdown"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When working with web content, converting HTML files to plain text or Markdown is a common requirement for content extraction, analysis, and preservation. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Content analysis&lt;/strong&gt;: Extract clean text from HTML for natural language processing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data mining&lt;/strong&gt;: Strip formatting to focus on the actual content&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Offline reading&lt;/strong&gt;: Convert web pages to readable formats for e-readers or offline consumption&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content migration&lt;/strong&gt;: Move content between different CMS platforms&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SEO analysis&lt;/strong&gt;: Extract headings, content structure, and text for optimization&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Archive creation&lt;/strong&gt;: Store web content in more compact, preservation-friendly formats&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;: Convert content to formats that work better with screen readers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This tutorial covers both converting existing HTML files and combining web crawling with HTML-to-text conversion in a single workflow &amp;ndash; all using the command line.&lt;/p&gt;</description></item><item><title/><link>/2026-05/convert-pdfs-to-markdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/convert-pdfs-to-markdown/</guid><description>&lt;h2 id="converting-pdfs-to-markdown"&gt;Converting PDFs to Markdown&lt;a class="anchor" href="#converting-pdfs-to-markdown"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;PDF documents are ubiquitous in academic, business, and technical contexts, but extracting and repurposing their content can be challenging. This tutorial explores various command-line tools for converting PDFs to Markdown format, with a focus on preserving structure and formatting suitable for different use cases, including preparation for Large Language Models (LLMs).&lt;/p&gt;
&lt;p&gt;Use Cases:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LLM training and fine-tuning&lt;/strong&gt;: Create clean text data from PDFs for AI model training&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Knowledge base creation&lt;/strong&gt;: Transform PDFs into searchable, editable markdown documents&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content repurposing&lt;/strong&gt;: Convert academic papers and reports for web publication&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data extraction&lt;/strong&gt;: Pull structured content from PDF documents for analysis&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;: Convert PDFs to more accessible formats for screen readers&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Citation and reference management&lt;/strong&gt;: Extract bibliographic information from academic papers&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Documentation conversion&lt;/strong&gt;: Transform technical PDFs into maintainable documentation&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="pymupdf4llm"&gt;PyMuPDF4LLM&lt;a class="anchor" href="#pymupdf4llm"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://pymupdf.readthedocs.io/en/latest/pymupdf4llm/"&gt;PyMuPDF4LLM&lt;/a&gt; is a specialized component of the PyMuPDF library that generates Markdown specifically formatted for Large Language Models. It produces high-quality markdown with good preservation of document structure. It&amp;rsquo;s specifically optimized for producing text that works well with LLMs, removing irrelevant formatting while preserving semantic structure. Requires PyTorch, which adds dependencies but enables more advanced processing capabilities.&lt;/p&gt;</description></item><item><title/><link>/2026-05/correlation-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/correlation-with-excel/</guid><description>&lt;h2 id="correlation-with-excel"&gt;Correlation with Excel&lt;a class="anchor" href="#correlation-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/lXHCyhO7DmY"&gt;&lt;img src="https://i.ytimg.com/vi_webp/lXHCyhO7DmY/sddefault.webp" alt="Correlation with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn to calculate and interpret correlations using Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Enabling the Data Analysis Tool Pack&lt;/strong&gt;: Steps to enable the Excel data analysis tool pack.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Correlation Analysis&lt;/strong&gt;: Understanding statistical association between variables.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating a Correlation Matrix&lt;/strong&gt;: Steps to generate and interpret a correlation matrix.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scatterplots and Trendlines&lt;/strong&gt;: Plotting data and adding trend lines to visualize correlations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Analyzing Results&lt;/strong&gt;: Comparing correlation coefficients and understanding their implications.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Insights and Further Analysis&lt;/strong&gt;: Interpreting scatterplots and planning further analysis for deeper insights.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/cors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/cors/</guid><description>&lt;h2 id="cors-cross-origin-resource-sharing"&gt;CORS: Cross-Origin Resource Sharing&lt;a class="anchor" href="#cors-cross-origin-resource-sharing"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;CORS (Cross-Origin Resource Sharing) is a security mechanism that controls how web browsers handle requests between different origins (domains, protocols, or ports). Data scientists need CORS for APIs serving data or analysis to a browser on a different domain.&lt;/p&gt;
&lt;p&gt;Watch this practical explanation of CORS (3 min):&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/4KHiSt0oLJ0"&gt;&lt;img src="https://i.ytimg.com/vi_webp/4KHiSt0oLJ0/sddefault.webp" alt="CORS in 100 Seconds" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Key CORS concepts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Same-Origin Policy&lt;/strong&gt;: Browsers block requests between different origins by default&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CORS Headers&lt;/strong&gt;: Server responses must include specific headers to allow cross-origin requests&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preflight Requests&lt;/strong&gt;: Browsers send OPTIONS requests to check if the actual request is allowed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Credentials&lt;/strong&gt;: Special handling required for requests with cookies or authentication&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you&amp;rsquo;re exposing your API with a GET request publicly, the only thing you need to do is set the HTTP header &lt;code&gt;Access-Control-Allow-Origin: *&lt;/code&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/crawling-cli/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/crawling-cli/</guid><description>&lt;h2 id="crawling-with-the-cli"&gt;Crawling with the CLI&lt;a class="anchor" href="#crawling-with-the-cli"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Since websites are a common source of data, we often download entire websites (crawling) and then process them offline.&lt;/p&gt;
&lt;p&gt;Web crawling is essential in many data-driven scenarios:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data mining and analysis&lt;/strong&gt;: Gathering structured data from multiple pages for market research, competitive analysis, or academic research&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content archiving&lt;/strong&gt;: Creating offline copies of websites for preservation or backup purposes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SEO analysis&lt;/strong&gt;: Analyzing site structure, metadata, and content to improve search rankings&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Legal compliance&lt;/strong&gt;: Capturing website content for regulatory or compliance documentation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Website migration&lt;/strong&gt;: Creating a complete copy before moving to a new platform or design&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Offline access&lt;/strong&gt;: Downloading educational resources, documentation, or reference materials for use without internet connection&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The most commonly used tool for fetching websites is &lt;a href="https://www.gnu.org/software/wget/"&gt;&lt;code&gt;wget&lt;/code&gt;&lt;/a&gt;. It is pre-installed in many UNIX distributions and easy to install.&lt;/p&gt;</description></item><item><title/><link>/2026-05/css-selectors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/css-selectors/</guid><description>&lt;h2 id="css-selectors"&gt;CSS Selectors&lt;a class="anchor" href="#css-selectors"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;CSS selectors are patterns used to select and style HTML elements on a web page. They are fundamental to web development and data scraping, allowing you to precisely target elements for styling or extraction.&lt;/p&gt;
&lt;p&gt;For data scientists, understanding CSS selectors is crucial when:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Web scraping with tools like Beautiful Soup or Scrapy&lt;/li&gt;
&lt;li&gt;Selecting elements for browser automation with Selenium&lt;/li&gt;
&lt;li&gt;Styling data visualizations and web applications&lt;/li&gt;
&lt;li&gt;Debugging website issues using browser DevTools&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Watch this comprehensive introduction to CSS selectors (20 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-aggregation-in-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-aggregation-in-excel/</guid><description>&lt;h2 id="data-aggregation-in-excel"&gt;Data Aggregation in Excel&lt;a class="anchor" href="#data-aggregation-in-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/NkpT0dDU8Y4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/NkpT0dDU8Y4/sddefault.webp" alt="Data aggregation in Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn data aggregation and visualization techniques in Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Cleanup&lt;/strong&gt;: Remove empty columns and rows with missing values.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating Excel Tables&lt;/strong&gt;: Convert raw data into tables for easier manipulation and formula application.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Date Manipulation&lt;/strong&gt;: Extract week, month, and year from date columns using Excel functions (WEEKNUM, TEXT).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Color Scales&lt;/strong&gt;: Apply color scales to visualize clusters and trends in data over time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pivot Tables&lt;/strong&gt;: Create pivot tables to aggregate data by location and date, summarizing values weekly and monthly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sparklines&lt;/strong&gt;: Use sparklines to visualize trends within pivot tables, making data patterns more apparent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Bars&lt;/strong&gt;: Implement data bars for graphical illustrations of numerical columns, showing trends and waves.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-analysis-with-datasette/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-analysis-with-datasette/</guid><description>&lt;h1 id="data-analysis-with-datasette"&gt;Data Analysis with Datasette&lt;a class="anchor" href="#data-analysis-with-datasette"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/7kDFBnXaw-c"&gt;&lt;img src="https://i.ytimg.com/vi_webp/7kDFBnXaw-c/sddefault.webp" alt="Introduction to Datasette and sqlite-utils" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Datasette is an open-source tool for exploring and publishing data. Created by Simon Willison, it turns any SQLite database into an instant web interface with powerful exploration features, JSON APIs, and visualization capabilities—all without writing code.&lt;/p&gt;
&lt;p&gt;Unlike traditional database tools that require SQL knowledge upfront, Datasette provides an interactive interface for exploring data through faceting, filtering, and full-text search. It&amp;rsquo;s particularly powerful for data journalism, analysis workflows, and sharing datasets with non-technical audiences.&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-analysis-with-duckdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-analysis-with-duckdb/</guid><description>&lt;h2 id="data-analysis-with-duckdb"&gt;Data Analysis with DuckDB&lt;a class="anchor" href="#data-analysis-with-duckdb"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/4U0GqYrET5s"&gt;&lt;img src="https://i.ytimg.com/vi_webp/4U0GqYrET5s/sddefault.webp" alt="Data Analysis with DuckDB" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to perform data analysis using DuckDB and Pandas, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Parquet for Data Storage&lt;/strong&gt;: Understand why Parquet is a faster, more compact, and better-typed storage format compared to CSV, JSON, and SQLite.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DuckDB Setup&lt;/strong&gt;: Learn how to install and set up DuckDB, along with integrating it into a Jupyter notebook environment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;File Format Comparisons&lt;/strong&gt;: Compare file formats by speed and size, observing the performance difference between saving and loading data in CSV, JSON, SQLite, and Parquet.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Faster Queries with DuckDB&lt;/strong&gt;: Learn how DuckDB uses parallel processing, columnar storage, and on-disk operations to outperform Pandas in speed and memory efficiency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SQL Query Execution in DuckDB&lt;/strong&gt;: Run SQL queries directly on Parquet files and Pandas DataFrames to compute metrics such as the number of unique flight routes delayed by certain time intervals.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory Efficiency&lt;/strong&gt;: Understand how DuckDB performs analytics without loading entire datasets into memory, making it highly efficient for large-scale data analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mixing DuckDB and Pandas&lt;/strong&gt;: Learn to interleave DuckDB and Pandas operations, leveraging the strengths of both tools to perform complex queries like correlations and aggregations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ranking and Filtering Data&lt;/strong&gt;: Use SQL and Pandas to rank arrival delays by distance and extract key insights, such as the earliest flight arrival for each route.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Joining Data&lt;/strong&gt;: Create a cost analysis by joining datasets and calculating total costs of flight delays, demonstrating DuckDB&amp;rsquo;s speed in joining and aggregating large datasets.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-analysis-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-analysis-with-python/</guid><description>&lt;h2 id="data-analysis-with-python"&gt;Data Analysis with Python&lt;a class="anchor" href="#data-analysis-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/ZPfZH14FK90"&gt;&lt;img src="https://i.ytimg.com/vi_webp/ZPfZH14FK90/sddefault.webp" alt="Data Analysis with Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn practical data analysis techniques in Python using Pandas, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reading Parquet Files&lt;/strong&gt;: Utilize Pandas to read Parquet file formats for efficient data handling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dataframe Inspection&lt;/strong&gt;: Methods to preview and understand the structure of a dataset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pivot Tables&lt;/strong&gt;: Creating and interpreting pivot tables to summarize data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Percentage Calculations&lt;/strong&gt;: Normalize pivot table values to percentages for better insights.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Correlation Analysis&lt;/strong&gt;: Calculate and interpret correlation between variables, including significance testing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Statistical Significance&lt;/strong&gt;: Use statistical tests to determine the significance of observed correlations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Datetime Handling&lt;/strong&gt;: Extract and manipulate date and time information from datetime columns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Visualization&lt;/strong&gt;: Generate and customize heat maps to visualize data patterns effectively.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Leveraging AI&lt;/strong&gt;: Use ChatGPT to generate and refine analytical code, enhancing productivity and accuracy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-analysis-with-sql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-analysis-with-sql/</guid><description>&lt;h2 id="data-analysis-with-sql"&gt;Data Analysis with SQL&lt;a class="anchor" href="#data-analysis-with-sql"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/Xn3QkYrThbI"&gt;&lt;img src="https://i.ytimg.com/vi_webp/Xn3QkYrThbI/sddefault.webp" alt="Data Analysis with Databases" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to perform data analysis using SQL (via Python), covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Database Connection&lt;/strong&gt;: How to connect to a MySQL database using SQLAlchemy and Pandas.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SQL Queries&lt;/strong&gt;: Execute SQL queries directly from a Python environment to retrieve and analyze data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Counting Rows&lt;/strong&gt;: Use SQL to count the number of rows in a table.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;User Activity Analysis&lt;/strong&gt;: Query and identify top users by post count.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Post Concentration&lt;/strong&gt;: Determine if a small percentage of users contribute the majority of posts using SQL aggregation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Correlation Calculation&lt;/strong&gt;: Calculate the Pearson correlation coefficient between user attributes such as age and reputation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Regression Analysis&lt;/strong&gt;: Compute the regression slope to understand the relationship between views and reputation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Large Data&lt;/strong&gt;: Perform calculations on large datasets by fetching aggregated values from the database rather than entire datasets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Statistical Analysis in SQL&lt;/strong&gt;: Use SQL as a tool for statistical analysis, demonstrating its power beyond simple data retrieval.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Leveraging AI&lt;/strong&gt;: Use ChatGPT to generate SQL queries and Python code, enhancing productivity and accuracy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-analysis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-analysis/</guid><description>&lt;h1 id="data-analysis"&gt;Data analysis&lt;a class="anchor" href="#data-analysis"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://drive.google.com/file/d/1isjtxFa43CLIFlLpo8mwwQfBog9VlXYl/view"&gt;&lt;svg xmlns="http://www.w3.org/2000/svg" width="22" height="22" fill="currentColor" class="bi bi-broadcast-pin" viewBox="0 0 16 16"&gt;
&lt;path d="M3.05 3.05a7 7 0 0 0 0 9.9.5.5 0 0 1-.707.707 8 8 0 0 1 0-11.314.5.5 0 0 1 .707.707m2.122 2.122a4 4 0 0 0 0 5.656.5.5 0 1 1-.708.708 5 5 0 0 1 0-7.072.5.5 0 0 1 .708.708m5.656-.708a.5.5 0 0 1 .708 0 5 5 0 0 1 0 7.072.5.5 0 1 1-.708-.708 4 4 0 0 0 0-5.656.5.5 0 0 1 0-.708m2.122-2.12a.5.5 0 0 1 .707 0 8 8 0 0 1 0 11.313.5.5 0 0 1-.707-.707 7 7 0 0 0 0-9.9.5.5 0 0 1 0-.707zM6 8a2 2 0 1 1 2.5 1.937V15.5a.5.5 0 0 1-1 0V9.937A2 2 0 0 1 6 8"/&gt;
&lt;/svg&gt; &lt;span style="font-size: 24px; margin: 0 6px; vertical-align: bottom"&gt;Data Analysis: Introduction Podcast&lt;/span&gt;&lt;/a&gt; by &lt;a href="https://notebooklm.google.com/"&gt;NotebookLM&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-cleansing-in-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-cleansing-in-excel/</guid><description>&lt;h2 id="data-cleansing-in-excel"&gt;Data Cleansing in Excel&lt;a class="anchor" href="#data-cleansing-in-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/7du7xkqeu4s"&gt;&lt;img src="https://i.ytimg.com/vi_webp/7du7xkqeu4s/sddefault.webp" alt="Clean up data in Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn basic but essential data cleaning techniques in Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Find and Replace&lt;/strong&gt;: Use Ctrl+H to replace or remove specific terms (e.g., removing &amp;ldquo;[more]&amp;rdquo; from country names).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Changing Data Formats&lt;/strong&gt;: Convert columns from general to numerical format.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Removing Extra Spaces&lt;/strong&gt;: Use the TRIM function to clean up unnecessary spaces in text.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Identifying and Removing Blank Cells&lt;/strong&gt;: Highlight and delete entire rows with blank cells using the &amp;ldquo;Go To Special&amp;rdquo; function.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Removing Duplicates&lt;/strong&gt;: Use the &amp;ldquo;Remove Duplicates&amp;rdquo; feature to eliminate duplicate entries, demonstrated with country names.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-preparation-in-duckdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-preparation-in-duckdb/</guid><description>&lt;h2 id="data-preparation-in-duckdb"&gt;Data Preparation in DuckDB&lt;a class="anchor" href="#data-preparation-in-duckdb"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=Fg8n-O0dhrI&amp;amp;list=PLw2SS5iImhEThtiGNPiNenOr2tVvLj6H7&amp;amp;index=2"&gt;&lt;img src="https://i.ytimg.com/vi_webp/fZj6kTwXN1U/sddefault.webp" alt="DuckDB" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;DuckDB&amp;rsquo;s SQL engine can handle large files quickly. Below are common cleaning tasks using the DuckDB CLI.&lt;/p&gt;
&lt;h3 id="create-a-sample-dataset"&gt;Create a Sample Dataset&lt;a class="anchor" href="#create-a-sample-dataset"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Let&amp;rsquo;s create a sample dataset that mimics real business data patterns - incomplete customer records, time-series orders, and regional variations. Before working with messy production data, you need a controlled environment to test data cleaning techniques. This sample represents common e-commerce scenarios: missing customer info (20% of orders), seasonal patterns (15-day cycles), and geographic segmentation that drive business decisions like inventory placement and marketing campaigns.&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-preparation-in-the-editor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-preparation-in-the-editor/</guid><description>&lt;h2 id="data-preparation-in-the-editor"&gt;Data Preparation in the Editor&lt;a class="anchor" href="#data-preparation-in-the-editor"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/99lYu43L9uM"&gt;&lt;img src="https://i.ytimg.com/vi_webp/99lYu43L9uM/sddefault.webp" alt="Data preparation in the editor" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to use a text editor &lt;a href="https://code.visualstudio.com/"&gt;Visual Studio Code&lt;/a&gt; to process and clean data, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Format&lt;/strong&gt; JSON files&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find all&lt;/strong&gt; and multiple cursors to extract specific fields&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sort&lt;/strong&gt; lines&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Delete duplicate&lt;/strong&gt; lines&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Replace&lt;/strong&gt; text with multiple cursors&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://drive.google.com/file/d/1VEnKChf4i04iKsQfw0MwoJlfkOBGQ65B/view?usp=drive_link"&gt;City-wise product sales JSON&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/data-preparation-in-the-shell/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-preparation-in-the-shell/</guid><description>&lt;h2 id="data-preparation-in-the-shell"&gt;Data Preparation in the Shell&lt;a class="anchor" href="#data-preparation-in-the-shell"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/XEdy4WK70vU"&gt;&lt;img src="https://i.ytimg.com/vi_webp/XEdy4WK70vU/sddefault.webp" alt="Data preparation in the shell" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to use UNIX tools to process and clean data, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;curl&lt;/code&gt; (or &lt;code&gt;wget&lt;/code&gt;) to fetch data from websites.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gzip&lt;/code&gt; (or &lt;code&gt;xz&lt;/code&gt;) to compress and decompress files.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;wc&lt;/code&gt; to count lines, words, and characters in text.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;head&lt;/code&gt; and &lt;code&gt;tail&lt;/code&gt; to get the start and end of files.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;cut&lt;/code&gt; to extract specific columns from text.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uniq&lt;/code&gt; to de-duplicate lines.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sort&lt;/code&gt; to sort lines.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;grep&lt;/code&gt; to filter lines containing specific text.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sed&lt;/code&gt; to search and replace text.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;awk&lt;/code&gt; for more complex text processing.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://colab.research.google.com/drive/1KSFkQDK0v__XWaAaHKeQuIAwYV0dkTe8"&gt;Data preparation in the shell - Notebook&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-preparation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-preparation/</guid><description>&lt;h1 id="data-preparation"&gt;Data Preparation&lt;a class="anchor" href="#data-preparation"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Data preparation is crucial because raw data is rarely perfect.&lt;/p&gt;
&lt;p&gt;It often contains errors, inconsistencies, or missing values. For example, marks data may have &amp;lsquo;NA&amp;rsquo; or &amp;lsquo;absent&amp;rsquo; for non-attendees, which you need to handle.&lt;/p&gt;
&lt;p&gt;This section teaches you how to clean up data, convert it to different formats, aggregate it if required, and get a feel for the data before you analyze.&lt;/p&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-sourcing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-sourcing/</guid><description>&lt;h1 id="data-sourcing"&gt;Data Sourcing&lt;a class="anchor" href="#data-sourcing"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Before you do any kind of data science, you obviously have to get the data to be able to analyze it, visualize it, narrate it, and deploy it.
And what we are going to cover in this module is how you get the data.&lt;/p&gt;
&lt;p&gt;There are three ways you can get the data.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The first is you can &lt;strong&gt;download&lt;/strong&gt; the data. Either somebody gives you the data and says download it from here, or you are asked to download it from the internet because it&amp;rsquo;s a public data source. But that&amp;rsquo;s the first way—you download the data.&lt;/li&gt;
&lt;li&gt;The second way is you can &lt;strong&gt;query it&lt;/strong&gt; from somewhere. It may be on a database. It may be available through an API. It may be available through a library. But these are ways in which you can selectively query parts of the data and stitch it together.&lt;/li&gt;
&lt;li&gt;The third way is you have to &lt;strong&gt;scrape it&lt;/strong&gt;. It&amp;rsquo;s not directly available in a convenient form that you can query or download. But it is, in fact, on a web page. It&amp;rsquo;s available on a PDF file. It&amp;rsquo;s available in a Word document. It&amp;rsquo;s available on an Excel file. It&amp;rsquo;s kind of structured, but you will have to figure out that structure and extract it from there.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In this module, we will be looking at the tools that will help you either download from a data source or query from an API or from a database or from a library. And finally, how you can scrape from different sources.&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-storytelling-with-llms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-storytelling-with-llms/</guid><description>&lt;h1 id="data-storytelling-with-llms"&gt;Data Storytelling with LLMs&lt;a class="anchor" href="#data-storytelling-with-llms"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Large Language Models (LLMs) can help create compelling data stories by assisting at every step of the data-to-story value chain:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Data Engineering (scraping, cleaning)&lt;/li&gt;
&lt;li&gt;Data Analysis (modeling, insights)&lt;/li&gt;
&lt;li&gt;Data Visualization (charts, narratives)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://sanand0.github.io/talks/2025-06-27-data-design-by-dialogue/"&gt;Watch this talk (30m) on data storytelling with LLMs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/htc3LwVbPgI"&gt;&lt;img src="https://i.ytimg.com/vi_webp/htc3LwVbPgI/sddefault.webp" alt="Data Design by Dialogue (30m)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;a class="anchor" href="#prerequisites"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To follow this tutorial, you&amp;rsquo;ll need:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://chatgpt.com/"&gt;ChatGPT Plus&lt;/a&gt; ($20/month) - recommended for better models&lt;/li&gt;
&lt;li&gt;Basic understanding of data analysis concepts&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Optional but useful:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-storytelling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-storytelling/</guid><description>&lt;h1 id="data-storytelling"&gt;Data Storytelling&lt;a class="anchor" href="#data-storytelling"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/aF93i6zVVQg"&gt;&lt;img src="https://i.ytimg.com/vi_webp/aF93i6zVVQg/sddefault.webp" alt="Narrate a story" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="transcript"&gt;Transcript&lt;a class="anchor" href="#transcript"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Once you&amp;rsquo;ve analyzed your data and put it together into charts, you still want to communicate it in a way that people will remember, that people will share. &lt;strong&gt;And that&amp;rsquo;s where storytelling comes in.&lt;/strong&gt; What we&amp;rsquo;re going to look at in this module are ways of telling stories.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;There are many ways in which you can narrate a story, and one way is just to show the numbers.&lt;/strong&gt; This in itself can be powerful. What you see here, for example, is a correlation matrix which has conditional formatting on Excel that will help you understand how two securities are correlated. And it&amp;rsquo;s pretty easy to see that, for example, the Pakistani rupee is negatively correlated with many of the securities. &lt;strong&gt;Just the numbers can give people an understanding of what&amp;rsquo;s happening behind the scenes and convey a story.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-transformation-in-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-transformation-in-excel/</guid><description>&lt;h2 id="data-transformation-in-excel"&gt;Data Transformation in Excel&lt;a class="anchor" href="#data-transformation-in-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/gR2IY5Naja0"&gt;&lt;img src="https://i.ytimg.com/vi_webp/gR2IY5Naja0/sddefault.webp" alt="Data transformation in Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn data transformation techniques in Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Calculating Ratios&lt;/strong&gt;: Compute metro area to city area and metro population to city population ratios.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Using Pivot Tables&lt;/strong&gt;: Create pivot tables to aggregate data and identify outliers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filtering Data&lt;/strong&gt;: Apply filters in pivot tables to analyze specific subsets of data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Counting Data Occurrences&lt;/strong&gt;: Use pivot tables to count the frequency of specific entries.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating Charts&lt;/strong&gt;: Generate charts from pivot table data to visualize distributions and outliers.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-visualization-with-chatgpt/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-visualization-with-chatgpt/</guid><description>&lt;h1 id="data-visualization-with-chatgpt"&gt;Data Visualization with ChatGPT&lt;a class="anchor" href="#data-visualization-with-chatgpt"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;ChatGPT and other Large Language Models (LLMs) can help create compelling data visualizations by:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Finding and analyzing datasets&lt;/li&gt;
&lt;li&gt;Generating visualization code&lt;/li&gt;
&lt;li&gt;Improving visual design&lt;/li&gt;
&lt;li&gt;Creating data stories&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://sanand0.github.io/talks/2025-06-28-prompt-to-plot/"&gt;Watch this workshop (2h) on creating data visualizations with ChatGPT&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/SdDulR-1bBM"&gt;&lt;img src="https://i.ytimg.com/vi_webp/SdDulR-1bBM/sddefault.webp" alt="Prompt to Plot (2 hours)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="prerequisites"&gt;Prerequisites&lt;a class="anchor" href="#prerequisites"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To follow this tutorial, you&amp;rsquo;ll need:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://gemini.google.com/"&gt;Gemini&lt;/a&gt; (free) - good for processing images and video&lt;/li&gt;
&lt;li&gt;A &lt;a href="https://chatgpt.com/"&gt;ChatGPT&lt;/a&gt; Plus subscription ($20/month) - recommended for access to advanced models and coding capabilities&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/"&gt;GitHub account&lt;/a&gt; - for publishing visualizations&lt;/li&gt;
&lt;li&gt;Basic familiarity with HTML/CSS/JavaScript&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Other useful but optional tools include:&lt;/p&gt;</description></item><item><title/><link>/2026-05/data-visualization-with-seaborn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-visualization-with-seaborn/</guid><description>&lt;h2 id="data-visualization-with-seaborn"&gt;Data Visualization with Seaborn&lt;a class="anchor" href="#data-visualization-with-seaborn"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://seaborn.pydata.org/"&gt;Seaborn&lt;/a&gt; is a data visualization library for Python. It&amp;rsquo;s based on Matplotlib but a bit easier to use, and a bit prettier.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/6GUZXDef2U0"&gt;&lt;img src="https://i.ytimg.com/vi_webp/6GUZXDef2U0/sddefault.webp" alt="Seaborn Tutorial : Seaborn Full Course" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This video tutorial provides a comprehensive guide to Seaborn, a powerful data visualization library built on Matplotlib. You&amp;rsquo;ll learn how to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Understand Seaborn&amp;rsquo;s core purpose&lt;/strong&gt;: It&amp;rsquo;s a &lt;strong&gt;data visualization library built on Matplotlib&lt;/strong&gt; that simplifies plotting, often creating complex plots with &lt;strong&gt;just one line of code&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Install and set up Seaborn&lt;/strong&gt;: Learn how to install it using &lt;code&gt;pip&lt;/code&gt; or &lt;code&gt;conda&lt;/code&gt;, and how to import necessary libraries like &lt;strong&gt;NumPy, Pandas, and Matplotlib&lt;/strong&gt; for seamless integration.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Load and manage data&lt;/strong&gt;: Primarily use &lt;strong&gt;Seaborn&amp;rsquo;s built-in datasets&lt;/strong&gt; (e.g., &lt;code&gt;car_crashes&lt;/code&gt;, &lt;code&gt;tips&lt;/code&gt;, &lt;code&gt;flights&lt;/code&gt;, &lt;code&gt;iris&lt;/code&gt;, &lt;code&gt;attention&lt;/code&gt;) for practice, and understand how to load other file types via Pandas.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Create various distribution plots&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Distribution plots (&lt;code&gt;displot&lt;/code&gt;)&lt;/strong&gt;: Visualize &lt;strong&gt;univariate distributions&lt;/strong&gt; (distributions of a single variable), including &lt;strong&gt;histograms&lt;/strong&gt; and &lt;strong&gt;Kernel Density Estimation (KDE)&lt;/strong&gt; plots, and learn to define &lt;code&gt;bins&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Joint plots (&lt;code&gt;jointplot&lt;/code&gt;)&lt;/strong&gt;: &lt;strong&gt;Compare two distributions&lt;/strong&gt; by default as a &lt;strong&gt;scatter plot&lt;/strong&gt;, and generate &lt;strong&gt;regression lines&lt;/strong&gt; or show &lt;strong&gt;KDE&lt;/strong&gt; or &lt;strong&gt;hexagon distributions&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;KDE plots (&lt;code&gt;kdeplot&lt;/code&gt;)&lt;/strong&gt;: Create standalone plots for &lt;strong&gt;kernel density estimations&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pair plots (&lt;code&gt;pairplot&lt;/code&gt;)&lt;/strong&gt;: Plot &lt;strong&gt;relationships across all numerical values&lt;/strong&gt; in a DataFrame, showing &lt;strong&gt;histograms on the diagonal&lt;/strong&gt; and &lt;strong&gt;scatter plots elsewhere&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rug plots (&lt;code&gt;rugplot&lt;/code&gt;)&lt;/strong&gt;: Visualize single column data points as sticks, indicating where data is denser.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Master various categorical plots&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Bar plots (&lt;code&gt;barplot&lt;/code&gt;)&lt;/strong&gt;: Analyze distributions of &lt;strong&gt;categorical data against numerical data&lt;/strong&gt;, understanding &lt;strong&gt;variance bars&lt;/strong&gt; and how to change the &lt;strong&gt;aggregation estimator&lt;/strong&gt; (e.g., &lt;code&gt;median&lt;/code&gt;, &lt;code&gt;std&lt;/code&gt;, &lt;code&gt;cov&lt;/code&gt;) using NumPy functions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Count plots (&lt;code&gt;countplot&lt;/code&gt;)&lt;/strong&gt;: Simply &lt;strong&gt;count the number of occurrences&lt;/strong&gt; for categorical variables.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Box plots (&lt;code&gt;boxplot&lt;/code&gt;)&lt;/strong&gt;: &lt;strong&gt;Compare variables by showing quartiles&lt;/strong&gt;, median, standard deviation, whiskers, and outliers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Violin plots (&lt;code&gt;violinplot&lt;/code&gt;)&lt;/strong&gt;: A combination of &lt;strong&gt;box plots and KDE plots&lt;/strong&gt;, visualizing the &lt;strong&gt;density estimation of data points&lt;/strong&gt; and how to &lt;code&gt;split&lt;/code&gt; them for comparison.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Strip plots (&lt;code&gt;stripplot&lt;/code&gt;)&lt;/strong&gt;: Draw &lt;strong&gt;scatter plots with one categorical variable&lt;/strong&gt;, often used with box plots, and learn to use &lt;code&gt;jitter&lt;/code&gt; to spread points and &lt;code&gt;dodge&lt;/code&gt; to separate categories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Swarm plots (&lt;code&gt;swarmplot&lt;/code&gt;)&lt;/strong&gt;: Similar to strip plots but &lt;strong&gt;adjusts points to prevent overlap&lt;/strong&gt;, often &lt;strong&gt;layered on top of violin plots&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generate matrix plots for correlation and data patterns&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Heat maps (&lt;code&gt;heatmap&lt;/code&gt;)&lt;/strong&gt;: Visualize data in a &lt;strong&gt;matrix format&lt;/strong&gt;, requiring data to be prepared using &lt;strong&gt;correlation matrices (&lt;code&gt;.corr()&lt;/code&gt;) or pivot tables (&lt;code&gt;.pivot_table()&lt;/code&gt;)&lt;/strong&gt;, and how to add &lt;code&gt;annotations&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cluster maps (&lt;code&gt;clustermap&lt;/code&gt;)&lt;/strong&gt;: Create &lt;strong&gt;hierarchically clustered heat maps&lt;/strong&gt; that calculate distances and &lt;strong&gt;reposition data to find specific patterns and clusters&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Utilize powerful grid systems for complex visualizations&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Pair grids (&lt;code&gt;PairGrid&lt;/code&gt;)&lt;/strong&gt;: Gain &lt;strong&gt;specific control over plot placement&lt;/strong&gt; within a grid, allowing you to &lt;strong&gt;map different plot types&lt;/strong&gt; (e.g., scatter, histogram, KDE) to the upper, lower, or diagonal sections.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Facet grids (&lt;code&gt;FacetGrid&lt;/code&gt;)&lt;/strong&gt;: Print &lt;strong&gt;multiple plots in a grid&lt;/strong&gt;, defining columns and rows based on categorical data, and apply various styling options to individual subplots.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Create and customize regression plots&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Regression plots (&lt;code&gt;lmplot&lt;/code&gt;)&lt;/strong&gt;: Study relationships between numerical variables, customize markers, sizes, and colors, and &lt;strong&gt;separate data into columns or rows&lt;/strong&gt; based on other variables for multi-faceted analysis.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Apply extensive styling and customization options&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Set overall plot styles&lt;/strong&gt; (&lt;code&gt;sns.set_style&lt;/code&gt;) like &lt;code&gt;white&lt;/code&gt;, &lt;code&gt;darkgrid&lt;/code&gt;, &lt;code&gt;whitegrid&lt;/code&gt;, &lt;code&gt;dark&lt;/code&gt;, and &lt;code&gt;ticks&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adjust plot size&lt;/strong&gt; (&lt;code&gt;plt.figure(figsize=...)&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Change plot context&lt;/strong&gt; (&lt;code&gt;sns.set_context&lt;/code&gt;) for different uses: &lt;code&gt;paper&lt;/code&gt; (Jupyter), &lt;code&gt;talk&lt;/code&gt; (presentation), &lt;code&gt;poster&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Control font scale&lt;/strong&gt; (&lt;code&gt;font_scale&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Manage axis visibility&lt;/strong&gt; by turning &lt;code&gt;spines&lt;/code&gt; on or off.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Change plot color schemes using &lt;code&gt;palettes&lt;/code&gt;&lt;/strong&gt; and exploring Matplotlib&amp;rsquo;s &lt;code&gt;color maps&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reposition plot legends&lt;/strong&gt; for better readability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Customize marker symbols&lt;/strong&gt;, sizes, line widths, and edge colors in various plot types.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/data-visualization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/data-visualization/</guid><description>&lt;h1 id="data-visualization"&gt;Data visualization&lt;a class="anchor" href="#data-visualization"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/XkxRDql00UU"&gt;&lt;img src="https://i.ytimg.com/vi_webp/XkxRDql00UU/sddefault.webp" alt="Data visualization" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Effective visuals connect analysis to decisions. This module spotlights the tools that help you design narratives, build interactive experiences, and adapt outputs for every kind of audience.&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll practice:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Story-first planning&lt;/strong&gt; with data storytelling frameworks and LLM partners such as ChatGPT to refine messaging.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Portable presentations&lt;/strong&gt; spanning HTML decks in RevealJS, Markdown slides in Marp, and notebook-native sessions in Marimo.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Specialized builders&lt;/strong&gt; including RAWgraphs, Seaborn, Excel forecasting charts, Flourish animations, PowerPoint motion, and Kumu network diagrams.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Collaborative perspectives&lt;/strong&gt; through actor network visualizations and AI-assisted storytelling that keep stakeholders aligned.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Use these resources to turn complex analyses into clear, compelling stories wherever your viewers meet them.&lt;/p&gt;</description></item><item><title/><link>/2026-05/dbt/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/dbt/</guid><description>&lt;h2 id="data-transformation-with-dbt"&gt;Data Transformation with dbt&lt;a class="anchor" href="#data-transformation-with-dbt"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/5rNquRnNb4E"&gt;&lt;img src="https://i.ytimg.com/vi_webp/5rNquRnNb4E/sddefault.webp" alt="Data Transformation with dbt" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to transform data using dbt (data build tool), covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;dbt Fundamentals&lt;/strong&gt;: Understand what dbt is and how it brings software engineering practices to data transformation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Project Setup&lt;/strong&gt;: Learn how to initialize a dbt project, configure your warehouse connection, and structure your models&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Models and Materialization&lt;/strong&gt;: Create your first dbt models and understand different materialization strategies (view, table, incremental)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Testing and Documentation&lt;/strong&gt;: Implement data quality tests and auto-generate documentation for your data models&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Jinja Templating&lt;/strong&gt;: Use Jinja for dynamic SQL generation, making your transformations more maintainable and reusable&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;References and Dependencies&lt;/strong&gt;: Learn how to reference other models and manage model dependencies&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sources and Seeds&lt;/strong&gt;: Configure source data connections and manage static reference data&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Macros and Packages&lt;/strong&gt;: Create reusable macros and leverage community packages to extend functionality&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Incremental Models&lt;/strong&gt;: Optimize performance by only processing new or changed data&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deployment and Orchestration&lt;/strong&gt;: Set up dbt Cloud or integrate with Airflow for production deployment&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal dbt model example, &lt;code&gt;models/staging/stg_customers.sql&lt;/code&gt;:&lt;/p&gt;</description></item><item><title/><link>/2026-05/deployment-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/deployment-tools/</guid><description>&lt;h1 id="deployment-tools"&gt;Deployment Tools&lt;a class="anchor" href="#deployment-tools"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Publishing a project means moving assets off your laptop and into reliable environments. This module maps the tooling that packages content, automates delivery, and keeps services reachable for collaborators and users.&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll cover:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Content preparation&lt;/strong&gt; by compressing images and authoring Markdown so sites load fast and stay readable.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hosting options&lt;/strong&gt; from static GitHub Pages sites and Google Colab notebooks to Vercel serverless functions, Docker and Podman containers, and GitHub Codespaces devcontainers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automation and connectivity&lt;/strong&gt; using GitHub Actions CI/CD, secure ngrok tunnels, and REST APIs with correct CORS handling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Production fundamentals&lt;/strong&gt; including FastAPI backends, Google Auth integration, and local LLM endpoints with Ollama.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Together these tools let you launch prototypes quickly, iterate safely, and graduate data products into production-grade deployments.&lt;/p&gt;</description></item><item><title/><link>/2026-05/development-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/development-tools/</guid><description>&lt;h1 id="development-tools"&gt;Development Tools&lt;a class="anchor" href="#development-tools"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;NOTE&lt;/strong&gt;: The tools in this module are &lt;strong&gt;PRE-REQUISITES&lt;/strong&gt; for the course. You would have used most of these before. If most of this is new to you, please take this course later.&lt;/p&gt;
&lt;p&gt;Some tools are fundamental to data science because they are industry standards and widely used by data science professionals. Mastering these tools will align you with current best practices and making you more adaptable in a fast-evolving industry.&lt;/p&gt;</description></item><item><title/><link>/2026-05/devtools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/devtools/</guid><description>&lt;h2 id="browser-devtools"&gt;Browser: DevTools&lt;a class="anchor" href="#browser-devtools"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://developer.chrome.com/docs/devtools/overview/"&gt;Chrome DevTools&lt;/a&gt; is the de facto standard for web development and data analysis in the browser.
You&amp;rsquo;ll use this a lot when debugging and inspecting web pages.&lt;/p&gt;
&lt;p&gt;Here are the key features you&amp;rsquo;ll use most:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Elements Panel&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Inspect and modify HTML/CSS in real-time&lt;/li&gt;
&lt;li&gt;Copy CSS selectors for web scraping&lt;/li&gt;
&lt;li&gt;Debug layout issues with the Box Model&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;// Copy selector in Console
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;copy&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;$0&lt;/span&gt;); &lt;span style="color:#75715e"&gt;// Copies selector of selected element
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Console Panel&lt;/strong&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/docker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/docker/</guid><description>&lt;h2 id="containers-docker-podman"&gt;Containers: Docker, Podman&lt;a class="anchor" href="#containers-docker-podman"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.docker.com/"&gt;Docker&lt;/a&gt; and &lt;a href="https://podman.io/"&gt;Podman&lt;/a&gt; are containerization tools that package your application and its dependencies into a standardized unit for software development and deployment.&lt;/p&gt;
&lt;p&gt;Docker is the industry standard. Podman is compatible with Docker and has better security (and a slightly more open license). In this course, we recommend Podman but Docker works in the same way.&lt;/p&gt;
&lt;p&gt;Initialize the container engine:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;podman machine init
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;podman machine start&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Common Operations. (You can use &lt;code&gt;docker&lt;/code&gt; instead of &lt;code&gt;podman&lt;/code&gt; in the same way.)&lt;/p&gt;</description></item><item><title/><link>/2026-05/embeddings/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/embeddings/</guid><description>&lt;h2 id="embeddings-openai-and-local-models"&gt;Embeddings: OpenAI and Local Models&lt;a class="anchor" href="#embeddings-openai-and-local-models"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Embedding models convert text into a list of numbers. These are like a map of text in numerical form. Each number represents a feature, and similar texts will have numbers close to each other. So, if the numbers are similar, the text they represent mean something similar.&lt;/p&gt;
&lt;p&gt;This is useful because text similarity is important in many common problems:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Search&lt;/strong&gt;. Find similar documents to a query.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Classification&lt;/strong&gt;. Classify text into categories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clustering&lt;/strong&gt;. Group similar items into clusters.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anomaly Detection&lt;/strong&gt;. Find an unusual piece of text.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You can run embedding models locally or using an API. Local models are better for privacy and cost. APIs are better for scale and quality.&lt;/p&gt;</description></item><item><title/><link>/2026-05/extracting-audio-and-transcripts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/extracting-audio-and-transcripts/</guid><description>&lt;h2 id="extracting-audio-and-transcripts"&gt;Extracting Audio and Transcripts&lt;a class="anchor" href="#extracting-audio-and-transcripts"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h2 id="media-processing-ffmpeg"&gt;Media Processing: FFmpeg&lt;a class="anchor" href="#media-processing-ffmpeg"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://ffmpeg.org/"&gt;FFmpeg&lt;/a&gt; is the standard command-line tool for processing video and audio files. It&amp;rsquo;s essential for data scientists working with media files for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Extracting audio/video for machine learning&lt;/li&gt;
&lt;li&gt;Converting formats for web deployment&lt;/li&gt;
&lt;li&gt;Creating visualizations and presentations&lt;/li&gt;
&lt;li&gt;Processing large media datasets&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Basic Operations:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Basic conversion&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i input.mp4 output.avi
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Extract audio&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i input.mp4 -vn output.mp3
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Convert format without re-encoding&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i input.mkv -c copy output.mp4
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# High quality encoding (crf: 0-51, lower is better)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i input.mp4 -preset slower -crf &lt;span style="color:#ae81ff"&gt;18&lt;/span&gt; output.mp4&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Common Data Science Tasks:&lt;/p&gt;</description></item><item><title/><link>/2026-05/fastapi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/fastapi/</guid><description>&lt;h2 id="web-framework-fastapi"&gt;Web Framework: FastAPI&lt;a class="anchor" href="#web-framework-fastapi"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://fastapi.tiangolo.com/"&gt;FastAPI&lt;/a&gt; is a modern Python web framework for building APIs with automatic interactive documentation. It&amp;rsquo;s fast, easy to use, and designed for building production-ready REST APIs.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal FastAPI app, &lt;code&gt;app.py&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# /// script&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# requires-python = &amp;#34;&amp;gt;=3.11&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# dependencies = [&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# &amp;#34;fastapi&amp;#34;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# &amp;#34;uvicorn&amp;#34;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# ]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# ///&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; fastapi &lt;span style="color:#f92672"&gt;import&lt;/span&gt; FastAPI
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;app &lt;span style="color:#f92672"&gt;=&lt;/span&gt; FastAPI()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;@app.get&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;/&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;async&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;root&lt;/span&gt;():
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {&lt;span style="color:#e6db74"&gt;&amp;#34;message&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Hello!&amp;#34;&lt;/span&gt;}
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; __name__ &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;__main__&amp;#34;&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;import&lt;/span&gt; uvicorn
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; uvicorn&lt;span style="color:#f92672"&gt;.&lt;/span&gt;run(app, host&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;0.0.0.0&amp;#34;&lt;/span&gt;, port&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;8000&lt;/span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Run this with &lt;code&gt;uv run app.py&lt;/code&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/feedback-2025-05/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/feedback-2025-05/</guid><description>&lt;h1 id="may-2025-feedback"&gt;May 2025 Feedback&lt;a class="anchor" href="#may-2025-feedback"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Based on feedback from 33 students, via &lt;a href="https://chatgpt.com/share/68cba081-afc0-800c-9da3-75222e84a499"&gt;ChatGPT&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="insights"&gt;Insights&lt;a class="anchor" href="#insights"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;⭐ &lt;strong&gt;“Time-pressure” complainers actually rated ROE slightly &lt;em&gt;higher&lt;/em&gt;.&lt;/strong&gt;
Students who mentioned ROE time limits gave ROE &lt;strong&gt;2.61&lt;/strong&gt; vs &lt;strong&gt;2.33&lt;/strong&gt; for others (&lt;strong&gt;+0.28&lt;/strong&gt;). This suggests the most engaged cohort both &lt;em&gt;feels&lt;/em&gt; the time squeeze and still perceives value—hinting the exam is closer to a “desirable difficulty” than a pure frustration. (Concept supported by learning-science literature on “desirable difficulties”.) (&lt;a href="https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf" title="Creating Desirable Difficulties to Enhance Learning"&gt;bjorklab.psych.ucla.edu&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Project-2 is love-it or hate-it.&lt;/strong&gt;
Median is &lt;strong&gt;5/5&lt;/strong&gt; with many top scores, but variance is highest (σ≈1.54). Those flagging “difficulty/prereqs” rated P2 &lt;strong&gt;1.32 points lower&lt;/strong&gt; than others—so P2 is excellent for prepared learners and punishing for under-prepared ones (bimodal fit).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Calling out TA/Discourse correlates with &lt;em&gt;higher&lt;/em&gt; satisfaction.&lt;/strong&gt;
Mentions of TA/Discourse associate with slightly higher ratings across all four metrics (Δ overall &lt;strong&gt;+0.08&lt;/strong&gt;, P1 &lt;strong&gt;+0.15&lt;/strong&gt;, P2 &lt;strong&gt;+0.32&lt;/strong&gt;, ROE &lt;strong&gt;+0.05&lt;/strong&gt;). That’s counter to the intuition that “more TA talk = more complaints”; it reads as “the engaged students leaned on Discourse and benefited.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Students know it’s hard—and still rate learning high.&lt;/strong&gt;
Despite many “hard/fast” comments, overall learning is strong (~&lt;strong&gt;4.0/5&lt;/strong&gt;). That aligns with the course’s stated design (“ROE is hard”; “programming skills are a pre-requisite”), and suggests difficulty didn’t suppress perceived learning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“Alignment” is a minority complaint—but a leverage point.&lt;/strong&gt;
Only a few explicitly said “align projects/ROE with taught content,” yet those who did had lower ROE ratings (−0.57) and overall learning (−0.47). A small group may be disproportionately dissatisfied due to scope drift.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="interesting-observations"&gt;Interesting observations&lt;a class="anchor" href="#interesting-observations"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ROE dominates the pain map.&lt;/strong&gt; Time limit (45 min per the schedule), ambiguity, and “what exactly are we proving?” recur. A few ask for mock portals and clearer rubric.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;P2 teaches beyond comfort.&lt;/strong&gt; Many “pushes us beyond” / “real-world exposure” notes in “Best things,” while the “Worst” ask for more scaffolding. Classic split between stretch vs support.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ops gaps show up:&lt;/strong&gt; evaluation delays; release/schedule slips; a few network/portal reliability mentions; form UX complaints (Google Form structure).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tooling friction:&lt;/strong&gt; requests for dependable GPT/AI-Pipe access and deployment help; some struggled with “free versions” / rate limits. (AI-Pipe is in your ecosystem.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Students explicitly ask for “instant feedback”&lt;/strong&gt; on submissions (linting/auto-checks), not just grades—hinting appetite for continuous formative signals.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A minority asks for &lt;em&gt;tougher projects&lt;/em&gt; (and longer ROE).&lt;/strong&gt; The tails want even more challenge/time—a useful signal for differentiated tracks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Desire for predictable cadence&lt;/strong&gt; (earlier release, stick to schedule) appears in multiple corners even if counts are lower; these often co-occur with lower satisfaction.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="actions"&gt;Actions&lt;a class="anchor" href="#actions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Gate early with a hard readiness check + branching track.&lt;/strong&gt;
Make GA-1 (already linked) the &lt;em&gt;gate&lt;/em&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;If GA-1 &amp;lt; threshold:&lt;/strong&gt; auto-route to a “Foundation track” of P2 (same outcomes, more scaffolds, smaller surface area).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;If GA-1 ≥ threshold:&lt;/strong&gt; “Pro track” (current P2).
Rationale: removes P2 bimodality; keeps challenge for prepared students; respects “desirable difficulty” while preventing wipeouts. Implementation is trivial (routing by score in your existing exam infra).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ROE redesign: keep pressure, fix &lt;em&gt;friction&lt;/em&gt;.&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Keep 45-min “pressure-cooker” (it likely drives durable learning), &lt;em&gt;but&lt;/em&gt; reduce non-constructive friction:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Two 10-min micro-mocks&lt;/strong&gt; in the exact ROE environment the week prior (auto-graded, unlimited retries).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stable prompt rubric&lt;/strong&gt; with 2–3 canonical “what good looks like” exemplars.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Network-independent tasks&lt;/strong&gt; (bundle data; pre-cache assets) and a simple “connectivity hiccup” auto-extend of 3–5 min if the portal detects time-outs. (&lt;a href="https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf" title="Creating Desirable Difficulties to Enhance Learning"&gt;bjorklab.psych.ucla.edu&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TA/Discourse “SLA + Triage Bot.”&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Publish an SLA (e.g., TA initial touch &amp;lt;6 h, resolution &amp;lt;24 h).&lt;/li&gt;
&lt;li&gt;Add a &lt;strong&gt;Discourse clarifier bot&lt;/strong&gt; that: (i) de-dupes to canonical answers; (ii) proposes clearer wording for confusing prompts; (iii) tags to the right TA. (Start with a rules-only bot; graduate to LLM summaries later.) This leans into the positive correlation you’re already seeing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Instant feedback everywhere.&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Add pre-submission checks: schema validators, smoke tests, minimal unit tests for projects; “red/green” badges on submit.&lt;/li&gt;
&lt;li&gt;For P2, expose a &lt;em&gt;thin&lt;/em&gt; public checker (5–10 hidden tests retained for grading).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cadence discipline via CI.&lt;/strong&gt;
Turn “release on time” from a promise into automation: a release calendar in the repo, GitHub Actions that publish content at fixed timestamps, and an &lt;strong&gt;automated grading queue dashboard&lt;/strong&gt; (counts, average turnaround, SLA breaches).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tooling reliability: AI-Pipe + deployment helpers.&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Provide a one-click “fallback” API key via AI-Pipe with rate-limit hints and a local stub; publish &lt;strong&gt;two minimal deployment recipes&lt;/strong&gt; (Vercel/Cloudflare) with a pre-flight checklist. (Your repo already points to AI-Pipe.)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Micro-scaffolds for P2 only (don’t blunt the challenge).&lt;/strong&gt;
Ship &lt;em&gt;optional&lt;/em&gt; scaffolds: sample input/output pairs, data dictionaries, 2 starter snippets. These trim unproductive thrash while preserving the tough core.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rubric transparency + sample graded artifacts.&lt;/strong&gt;
Post 2 anonymized “A/B/C” exemplars with rubric mapping; students asked for “what exactly are we proving?” — this answers that.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Contrarian take:&lt;/p&gt;</description></item><item><title/><link>/2026-05/fine-tuning-gemini/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/fine-tuning-gemini/</guid><description>&lt;h2 id="fine-tuning-llms-gemini"&gt;Fine tuning LLMs: Gemini&lt;a class="anchor" href="#fine-tuning-llms-gemini"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This material will likely be added in the May 2025 term. It is not part of the Jan 2025 term.&lt;/p&gt;</description></item><item><title/><link>/2026-05/forecasting-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/forecasting-with-excel/</guid><description>&lt;h2 id="forecasting-with-excel"&gt;Forecasting with Excel&lt;a class="anchor" href="#forecasting-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/QrTimmxwZw4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/QrTimmxwZw4/sddefault.webp" alt="Forecasting with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/forecast-and-forecast-linear-functions-50ca49c9-7b40-4892-94e4-7ad38bbeda99"&gt;FORECAST reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/forecast-ets-function-15389b8b-677e-4fbd-bd95-21d464333f41"&gt;FORECAST.ETS reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.google.com/spreadsheets/d/1iMFVPh8q9KgnfLwBeBMmX1GaFabP02FK/view"&gt;Height-weight dataset&lt;/a&gt; from &lt;a href="https://www.kaggle.com/datasets/burnoutminer/heights-and-weights-dataset"&gt;Kaggle&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.google.com/spreadsheets/d/1w2R0fHdLG5ZGW-papaK7wzWq_-WDArKC/view"&gt;Traffic dataset&lt;/a&gt; from &lt;a href="https://www.kaggle.com/datasets/fedesoriano/traffic-prediction-dataset"&gt;Kaggle&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/function-calling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/function-calling/</guid><description>&lt;h2 id="function-calling-with-openai"&gt;Function Calling with OpenAI&lt;a class="anchor" href="#function-calling-with-openai"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://platform.openai.com/docs/guides/function-calling"&gt;Function Calling&lt;/a&gt; allows Large Language Models to convert natural language into structured function calls. This is perfect for building chatbots and AI assistants that need to interact with your backend systems.&lt;/p&gt;
&lt;p&gt;OpenAI supports &lt;a href="https://platform.openai.com/docs/guides/function-calling"&gt;Function Calling&lt;/a&gt; &amp;ndash; a way for LLMs to suggest what functions to call and how.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/aqdWSYWC_LI"&gt;&lt;img src="https://i.ytimg.com/vi_webp/aqdWSYWC_LI/sddefault.webp" alt="OpenAI Function Calling - Full Beginner Tutorial" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal example using Python and OpenAI&amp;rsquo;s function calling that identifies the weather in a given location.&lt;/p&gt;</description></item><item><title/><link>/2026-05/geospatial-analysis-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/geospatial-analysis-with-excel/</guid><description>&lt;h2 id="geospatial-analysis-with-excel"&gt;Geospatial Analysis with Excel&lt;a class="anchor" href="#geospatial-analysis-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/49LjxNvxyVs"&gt;&lt;img src="https://i.ytimg.com/vi_webp/49LjxNvxyVs/sddefault.webp" alt="Geospatial analysis with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to create a data-driven story about coffee shop coverage in Manhattan, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Collection&lt;/strong&gt;: Collect and scrape data for coffee shop locations and census population from various sources.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Processing&lt;/strong&gt;: Use Python libraries like geopandas for merging population data with geographic maps.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Map Creation&lt;/strong&gt;: Generate coverage maps using tools like QGIS and Excel to visualize coffee shop distribution and population impact.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visualization&lt;/strong&gt;: Create physical, Power BI, and video visualizations to present the data effectively.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storytelling&lt;/strong&gt;: Craft a narrative around coffee shop competition, including strategic insights and potential market changes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links that explain how the video was made:&lt;/p&gt;</description></item><item><title/><link>/2026-05/geospatial-analysis-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/geospatial-analysis-with-python/</guid><description>&lt;h2 id="geospatial-analysis-with-python"&gt;Geospatial Analysis with Python&lt;a class="anchor" href="#geospatial-analysis-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/m_qayAJt-yE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/m_qayAJt-yE/sddefault.webp" alt="Geospatial analysis with Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to perform geospatial analysis for location-based decision making, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Distance Calculation&lt;/strong&gt;: Compute distances between various store locations and a reference point, such as the Empire State Building.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Visualization&lt;/strong&gt;: Visualize store locations on a map using Python libraries like Folium.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Store Density Analysis&lt;/strong&gt;: Determine the number of stores within a specified radius.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Proximity Analysis&lt;/strong&gt;: Identify the closest and farthest stores from a specific location.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decision Making&lt;/strong&gt;: Use geospatial data to assess whether opening a new store is feasible based on existing store distribution.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/geospatial-analysis-with-qgis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/geospatial-analysis-with-qgis/</guid><description>&lt;h2 id="geospatial-analysis-with-qgis"&gt;Geospatial Analysis with QGIS&lt;a class="anchor" href="#geospatial-analysis-with-qgis"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/tJhehs0o-ik"&gt;&lt;img src="https://i.ytimg.com/vi_webp/tJhehs0o-ik/sddefault.webp" alt="Geospatial analysis with QGIS" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to use QGIS for geographic data processing, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Shapefiles and KML Files&lt;/strong&gt;: Create and manage shapefiles and KML files for storing and analyzing geographic information.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downloading QGIS&lt;/strong&gt;: Install QGIS on different operating systems and familiarize yourself with its interface.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Geospatial Data&lt;/strong&gt;: Access and utilize shapefiles from sources like Diva-GIS and integrate them into QGIS projects.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating Custom Shapefiles&lt;/strong&gt;: Learn how to create custom shapefiles when existing ones are unavailable, including creating a shapefile for South Sudan.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Editing and Visualization&lt;/strong&gt;: Use QGIS tools to edit shapefiles, add attributes, and visualize geographic data with various styling and labeling options.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exporting Data&lt;/strong&gt;: Export shapefiles or KML files for use in other applications, such as Google Earth.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/git/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/git/</guid><description>&lt;h2 id="version-control-git-github"&gt;Version Control: Git, GitHub&lt;a class="anchor" href="#version-control-git-github"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://git-scm.com/"&gt;Git&lt;/a&gt; is the de facto standard for version control of software (and sometimes, data as well). It&amp;rsquo;s a system that keeps track of changes you make to files and folders. It allows you to revert to a previous state, compare changes, etc. It&amp;rsquo;s a central tool in any developer&amp;rsquo;s workflow.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/"&gt;GitHub&lt;/a&gt; is the most popular hosting service for Git repositories. It&amp;rsquo;s a website that shows your code, allows you to collaborate with others, and provides many useful tools for developers.&lt;/p&gt;</description></item><item><title/><link>/2026-05/github-actions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/github-actions/</guid><description>&lt;h2 id="cicd-github-actions"&gt;CI/CD: GitHub Actions&lt;a class="anchor" href="#cicd-github-actions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/features/actions"&gt;GitHub Actions&lt;/a&gt; is a powerful automation platform built into GitHub. It helps automate your development workflow - running tests, deploying applications, updating datasets, retraining models, etc.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Understand the basics of &lt;a href="https://docs.github.com/en/actions/writing-workflows/quickstart"&gt;YAML configuration files&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Explore the &lt;a href="https://github.com/marketplace?type=actions"&gt;pre-built actions from the marketplace&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;How to &lt;a href="https://docs.github.com/en/actions/security-for-github-actions/security-guides/using-secrets-in-github-actions"&gt;handle secrets securely&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/actions/writing-workflows/choosing-when-your-workflow-runs/triggering-a-workflow"&gt;Triggering a workflow&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Staying within the &lt;a href="https://docs.github.com/en/billing/managing-billing-for-your-products/managing-billing-for-github-actions/about-billing-for-github-actions"&gt;free tier limits&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/actions/writing-workflows/choosing-what-your-workflow-does/caching-dependencies-to-speed-up-workflows"&gt;Caching dependencies to speed up workflows&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here is a sample &lt;code&gt;.github/workflows/iss-location.yml&lt;/code&gt; that runs daily, appends the International Space Station location data into &lt;code&gt;iss-location.jsonl&lt;/code&gt;, and commits it to the repository.&lt;/p&gt;</description></item><item><title/><link>/2026-05/github-codespaces/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/github-codespaces/</guid><description>&lt;h2 id="ide-github-codespaces"&gt;IDE: GitHub Codespaces&lt;a class="anchor" href="#ide-github-codespaces"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/features/codespaces"&gt;GitHub Codespaces&lt;/a&gt; is a cloud-hosted development environment built right into GitHub that gets you coding faster with pre-configured containers, adjustable compute power, and seamless integration with workflows like Actions and Copilot.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why Codespaces helps&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reproducible onboarding&lt;/strong&gt;: Say goodbye to “works on my machine” woes—everyone uses the same setup for assignments or demos.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anywhere access&lt;/strong&gt;: Jump back into your project from a laptop, tablet, or phone without having to reinstall anything.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rapid experimentation &amp;amp; debugging&lt;/strong&gt;: Spin up short-lived environments on any branch, commit, or PR to isolate bugs or test features, or keep longer-lived codespaces for big projects.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=-tQ2nxjqP6o"&gt;&lt;img src="https://i.ytimg.com/vi_webp/-tQ2nxjqP6o/sddefault.webp" alt="Introduction to GitHub Codespaces (5 min)" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/github-copilot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/github-copilot/</guid><description>&lt;h2 id="ai-editor-github-copilot"&gt;AI Editor: GitHub Copilot&lt;a class="anchor" href="#ai-editor-github-copilot"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;AI Code Editors like &lt;a href="https://github.com/features/copilot"&gt;GitHub Copilot&lt;/a&gt;, &lt;a href="https://www.cursor.com/"&gt;Cursor&lt;/a&gt;, &lt;a href="http://windsurf.com/"&gt;Windsurf&lt;/a&gt;, &lt;a href="https://roocode.com/"&gt;Roo Code&lt;/a&gt;, &lt;a href="https://cline.bot/"&gt;Cline&lt;/a&gt;, &lt;a href="https://www.continue.dev/"&gt;Continue.dev&lt;/a&gt;, etc. use LLMs to help you write code faster.&lt;/p&gt;
&lt;p&gt;Most are built on top of &lt;a href="/2026-05/vscode/"&gt;VS Code&lt;/a&gt;. These are now a standard tool in every developer&amp;rsquo;s toolkit.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/features/copilot"&gt;GitHub Copilot&lt;/a&gt; is &lt;a href="https://github.com/features/copilot/plans"&gt;free&lt;/a&gt; (as of May 2025) for 2,000 completions and 50 chats.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/n0NlxUyA7FI"&gt;&lt;img src="https://i.ytimg.com/vi_webp/n0NlxUyA7FI/sddefault.webp" alt="Getting started with GitHub Copilot | Tutorial (11 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You should learn about:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/enterprise-cloud@latest/copilot/using-github-copilot/using-github-copilot-code-suggestions-in-your-editor"&gt;Code Suggestions&lt;/a&gt;, which is a basic feature.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/copilot/github-copilot-chat/using-github-copilot-chat-in-your-ide"&gt;Using Chat&lt;/a&gt;, which lets you code in natural language.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/copilot/using-github-copilot/ai-models/changing-the-ai-model-for-copilot-chat"&gt;Changing the chat model&lt;/a&gt;. The free version includes Claude 3.5 Sonnet, a good coding model.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/copilot/copilot-chat-cookbook"&gt;Prompts&lt;/a&gt; to understand how people use AI code editors.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/github-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/github-pages/</guid><description>&lt;h2 id="static-hosting-github-pages"&gt;Static hosting: GitHub Pages&lt;a class="anchor" href="#static-hosting-github-pages"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://pages.github.com/"&gt;GitHub Pages&lt;/a&gt; is a free hosting service that turns your GitHub repository directly into a static website whenever you push it. This is useful for sharing analysis results, data science portfolios, project documentation, and more.&lt;/p&gt;
&lt;p&gt;Common Operations:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Create a new GitHub repo&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mkdir my-site
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd my-site
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git init
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Add your static content&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;echo &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;lt;h1&amp;gt;My Site&amp;lt;/h1&amp;gt;&amp;#34;&lt;/span&gt; &amp;gt; index.html
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Push to GitHub&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git add .
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git commit -m &lt;span style="color:#e6db74"&gt;&amp;#34;feat(pages): initial commit&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git push origin main
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Enable GitHub Pages from the main branch on the repo settings page&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Best Practices:&lt;/p&gt;</description></item><item><title/><link>/2026-05/google-auth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/google-auth/</guid><description>&lt;h2 id="google-authentication-with-fastapi"&gt;Google Authentication with FastAPI&lt;a class="anchor" href="#google-authentication-with-fastapi"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Secure your API endpoints using Google ID tokens to restrict access to specific email addresses.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/4ExQYRCwbzw"&gt;&lt;img src="https://i.ytimg.com/vi_webp/4ExQYRCwbzw/sddefault.webp" alt="🔥 Python FastAPI Google Login Tutorial | OAuth2 Authentication (19 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Google Auth is the most commonly implemented single sign-on mechanism because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It&amp;rsquo;s popular and user-friendly. Users can log in with their existing Google accounts.&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s secure: Google supports OAuth2 and OpenID Connect to handle authentication.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&amp;rsquo;s how you build a FastAPI app that identifies the user.&lt;/p&gt;</description></item><item><title/><link>/2026-05/google-charts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/google-charts/</guid><description>&lt;h2 id="google-charts"&gt;Google Charts&lt;a class="anchor" href="#google-charts"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://youtu.be/-DQP4fpmJpc"&gt;Google Chart 1&lt;/a&gt;: How To Create Chart Or Graph On HTML CSS Website | Google Charts Tutorial&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/6tx58ZZGxrI"&gt;Google Chart 2&lt;/a&gt;: Using Google Charts with googleVis&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/google-data-studio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/google-data-studio/</guid><description>&lt;h2 id="google-data-studio"&gt;Google Data Studio&lt;a class="anchor" href="#google-data-studio"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://youtu.be/1qGsjmmHiu8"&gt;Google Data Studio&lt;/a&gt;: Google Data Studio Tutorial for Beginners🔥&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/hosting-llms-runpod/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/hosting-llms-runpod/</guid><description>&lt;h2 id="hosting-llms-on-runpod"&gt;Hosting LLMs on Runpod&lt;a class="anchor" href="#hosting-llms-on-runpod"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This material will likely be added in the May 2025 term. It is not part of the Jan 2025 term.&lt;/p&gt;</description></item><item><title/><link>/2026-05/http-requests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/http-requests/</guid><description>&lt;h2 id="http-requests-curl-wget-httpie-postman"&gt;HTTP Requests: curl, wget, HTTPie, Postman&lt;a class="anchor" href="#http-requests-curl-wget-httpie-postman"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Making HTTP requests is essential for interacting with web APIs, downloading data, and testing web services. Whether you&amp;rsquo;re fetching data from a public API, testing your own REST endpoints, or automating data collection, these tools make HTTP interactions simple and efficient.&lt;/p&gt;
&lt;p&gt;This guide covers four popular approaches: &lt;code&gt;curl&lt;/code&gt; (the universal standard), &lt;code&gt;wget&lt;/code&gt; (for downloads), &lt;code&gt;HTTPie&lt;/code&gt; (user-friendly alternative), and Postman (GUI-based testing).&lt;/p&gt;
&lt;p&gt;Watch these tutorials to understand HTTP requests and API testing (60 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/huggingface-spaces/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/huggingface-spaces/</guid><description>&lt;h2 id="container-hosting-hugging-face-spaces-with-docker"&gt;Container hosting: Hugging Face Spaces with Docker&lt;a class="anchor" href="#container-hosting-hugging-face-spaces-with-docker"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Container platforms let you deploy applications in isolated, portable environments that include all dependencies. They&amp;rsquo;re perfect for ML applications that need &lt;em&gt;custom environments, specific packages, or persistent state&lt;/em&gt;. Here are some common real-life uses:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A text summarization API that processes documents using custom NLP models (runs for 10-30 seconds per request)&lt;/li&gt;
&lt;li&gt;An image classification tool that requires specific computer vision libraries (runs for 5-15 seconds per upload)&lt;/li&gt;
&lt;li&gt;A chatbot with custom fine-tuned models that need GPU acceleration (runs for 2-5 seconds per message)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unlike serverless functions, containers can maintain state, install any packages, write to the filesystem, and run background processes. They provide more control over the environment but require more configuration.&lt;/p&gt;</description></item><item><title/><link>/2026-05/hybrid-rag-typesense/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/hybrid-rag-typesense/</guid><description>&lt;h2 id="hybrid-retrieval-augmented-generation-hybrid-rag-with-typesense"&gt;Hybrid Retrieval Augmented Generation (Hybrid RAG) with TypeSense&lt;a class="anchor" href="#hybrid-retrieval-augmented-generation-hybrid-rag-with-typesense"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Hybrid RAG combines semantic (vector) search with traditional keyword search to improve retrieval accuracy and relevance. By mixing exact text matches with embedding-based similarity, you get the best of both worlds: precision when keywords are present, and semantic recall when phrasing varies. &lt;a href="https://typesense.org/"&gt;TypeSense&lt;/a&gt; makes this easy with built-in hybrid search and automatic embedding generation.&lt;/p&gt;
&lt;p&gt;Below is a fully self-contained Hybrid RAG tutorial using TypeSense, Python, and the command line.&lt;/p&gt;</description></item><item><title/><link>/2026-05/image-compression/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/image-compression/</guid><description>&lt;h2 id="images-compression"&gt;Images: Compression&lt;a class="anchor" href="#images-compression"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Image compression is essential when deploying apps. Often, pages have dozens of images. Image analysis runs over thousands of images. The cost of storage and bandwidth can grow over time.&lt;/p&gt;
&lt;p&gt;Here are things you should know when you&amp;rsquo;re compressing images:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Image dimensions&lt;/strong&gt; are the width and height of the image in pixels. This impacts image size a lot&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lossless&lt;/strong&gt; compression (PNG, WebP) preserves exact data&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lossy&lt;/strong&gt; compression (JPEG, WebP) removes some data for smaller files&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vector&lt;/strong&gt; formats (SVG) scale without quality loss&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;WebP&lt;/strong&gt; is the modern standard, supporting both lossy and lossless&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&amp;rsquo;s a rule of thumb you can use as of 2025.&lt;/p&gt;</description></item><item><title/><link>/2026-05/json/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/json/</guid><description>&lt;h2 id="json"&gt;JSON&lt;a class="anchor" href="#json"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;JSON (JavaScript Object Notation) is the de facto standard format for data exchange on the web and APIs. Its human-readable format and widespread support make it essential for data scientists working with web services, APIs, and configuration files.&lt;/p&gt;
&lt;p&gt;For data scientists, JSON is essential when:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Working with REST APIs and web services&lt;/li&gt;
&lt;li&gt;Storing configuration files and metadata&lt;/li&gt;
&lt;li&gt;Parsing semi-structured data from databases like MongoDB&lt;/li&gt;
&lt;li&gt;Creating data visualization specifications (e.g., Vega-Lite)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Watch this comprehensive introduction to JSON (15 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/all-lab-assignments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/all-lab-assignments/</guid><description>&lt;h1 id="all-lab-assignments"&gt;All Lab Assignments&lt;a class="anchor" href="#all-lab-assignments"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="week-01--development-environment--tooling"&gt;Week 01 — Development Environment &amp;amp; Tooling&lt;a class="anchor" href="#week-01--development-environment--tooling"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-1/01-publish-python-library-pypi-uv/"&gt;Publish a Python library to PyPI using UV&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-1/02-burpsuite-traffic-debugging/"&gt;Web Traffic Debugging with Burp Suite&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-02--deployment--api-engineering"&gt;Week 02 — Deployment &amp;amp; API Engineering&lt;a class="anchor" href="#week-02--deployment--api-engineering"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-2/01-llm-deployment/"&gt;Deploying Private LLMs with vLLM &amp;amp; API Gateway&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-2/02-socket-chat/"&gt;WebSocket Chat with Redis &amp;amp; PostgreSQL&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-03--llm-engineering"&gt;Week 03 — LLM Engineering&lt;a class="anchor" href="#week-03--llm-engineering"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-3/youtube-subtitles-topics-json-pipeline/"&gt;YouTube → Subtitles → Topics → timestamps + JSON summary pipeline&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-3/cost-tracking-dashboard-langsmith/"&gt;Cost-tracking dashboard; compare prompt strategies via LangSmith&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-04--rag--hybrid-rag"&gt;Week 04 — RAG &amp;amp; Hybrid RAG&lt;a class="anchor" href="#week-04--rag--hybrid-rag"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CAPSTONE&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-4/capstone-bs-degree-chatbot/"&gt;BS Degree Chatbot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-4/ragas-evaluation-dashboard/"&gt;RAGAS evaluation dashboard (Naive vs Hybrid vs Contextual)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-05--agentic-ai"&gt;Week 05 — Agentic AI&lt;a class="anchor" href="#week-05--agentic-ai"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CAPSTONE&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-5/capstone-autonomous-research-agent/"&gt;Autonomous Research Agent&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-06--web-data-acquisition--osint"&gt;Week 06 — Web Data Acquisition &amp;amp; OSINT&lt;a class="anchor" href="#week-06--web-data-acquisition--osint"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CAPSTONE&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-6/capstone-job-posting-scraper-tracker/"&gt;Job Posting Scraper &amp;amp; Tracker&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CAPSTONE&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-6/capstone-ai-signature-detection-cropper/"&gt;AI Signature Detection &amp;amp; Cropper&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CAPSTONE&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-6/capstone-live-multilingual-travel-translator/"&gt;Live Multilingual Travel Translator&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-6/scheduled-scraper-github-actions-cron/"&gt;Scheduled scraper with GitHub Actions cron (DuckDB + Parquet)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-6/osint-open-source-dossier/"&gt;Open-Source Organisation Dossier (cited OSINT)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="week-07--cicd-security--cloud-infrastructure"&gt;Week 07 — CI/CD, Security &amp;amp; Cloud Infrastructure&lt;a class="anchor" href="#week-07--cicd-security--cloud-infrastructure"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-7/01-red-team-your-api-guardrails/"&gt;Red-team your own LLM API (attack → defend → regression suite)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lab&lt;/strong&gt;: &lt;a href="/2026-05/labs/week-7/02-full-cicd-cloud-run/"&gt;Full CI/CD: build → scan → Artifact Registry → deploy to Cloud Run&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/labs/capstone-projects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/capstone-projects/</guid><description>&lt;h1 id="capstone-projects"&gt;Capstone Projects&lt;a class="anchor" href="#capstone-projects"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;#&lt;/th&gt;
					&lt;th&gt;Project&lt;/th&gt;
					&lt;th&gt;Week&lt;/th&gt;
					&lt;th&gt;Description&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;C1&lt;/td&gt;
					&lt;td&gt;&lt;a href="/2026-05/labs/week-4/capstone-bs-degree-chatbot/"&gt;BS Degree Chatbot&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;Week 04&lt;/td&gt;
					&lt;td&gt;Hybrid RAG chatbot + RAGAS evaluation&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;C2&lt;/td&gt;
					&lt;td&gt;&lt;a href="/2026-05/labs/week-5/capstone-autonomous-research-agent/"&gt;Autonomous Research Agent&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;Week 05&lt;/td&gt;
					&lt;td&gt;LangChain daily research blog with memory, validation, and GitHub Actions&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;C3&lt;/td&gt;
					&lt;td&gt;&lt;a href="/2026-05/labs/week-6/capstone-job-posting-scraper-tracker/"&gt;Job Posting Scraper &amp;amp; Tracker&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;Week 06&lt;/td&gt;
					&lt;td&gt;Hidden API → validated schema → SQLite state → Parquet → DuckDB dashboard&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;C4&lt;/td&gt;
					&lt;td&gt;&lt;a href="/2026-05/labs/week-6/capstone-ai-signature-detection-cropper/"&gt;AI Signature Detection &amp;amp; Cropper&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;Week 06&lt;/td&gt;
					&lt;td&gt;Detect → crop → independent verifier, scored on precision/recall&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;C5&lt;/td&gt;
					&lt;td&gt;&lt;a href="/2026-05/labs/week-6/capstone-live-multilingual-travel-translator/"&gt;Live Multilingual Travel Translator&lt;/a&gt;&lt;/td&gt;
					&lt;td&gt;Week 06&lt;/td&gt;
					&lt;td&gt;STT → translate → TTS pipeline&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/labs/week-1/01-publish-python-library-pypi-uv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-1/01-publish-python-library-pypi-uv/</guid><description>&lt;h1 id="lab-11--publish-a-python-library-to-pypi-using-uv"&gt;Lab 1.1 — Publish a Python Library to PyPI using UV&lt;a class="anchor" href="#lab-11--publish-a-python-library-to-pypi-using-uv"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;What you&amp;rsquo;ll build&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;A real Python package, installable by the world via &lt;code&gt;pip install your-package-name&lt;/code&gt;, published to PyPI from GitHub Actions using &lt;strong&gt;Trusted Publishing&lt;/strong&gt; (no API tokens, no secrets in your repo). We&amp;rsquo;ll use &lt;strong&gt;UV&lt;/strong&gt; for every step.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;strong&gt;Time:&lt;/strong&gt; 60–90 minutes.
&lt;strong&gt;Difficulty:&lt;/strong&gt; ⭐⭐☆☆☆.
&lt;strong&gt;Ship:&lt;/strong&gt; your name on &lt;a href="https://pypi.org"&gt;pypi.org&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="what-the-finished-thing-looks-like"&gt;What the Finished Thing Looks Like&lt;a class="anchor" href="#what-the-finished-thing-looks-like"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;By the end:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;pip install tds-hello-&amp;lt;yourname&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python -c &lt;span style="color:#e6db74"&gt;&amp;#34;from tds_hello import greet; print(greet(&amp;#39;World&amp;#39;))&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Hello, World! — from tds-hello v0.1.0&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;And every &lt;code&gt;git tag v*&lt;/code&gt; push auto-publishes a new version.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-1/02-burpsuite-traffic-debugging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-1/02-burpsuite-traffic-debugging/</guid><description>&lt;h1 id="lab-12--web-traffic-debugging-with-burp-suite"&gt;Lab 1.2 — Web Traffic Debugging with Burp Suite&lt;a class="anchor" href="#lab-12--web-traffic-debugging-with-burp-suite"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;What you&amp;rsquo;ll build&lt;/strong&gt;
A local HTTP proxy laboratory. You will write a single-file Python script to host an interactive security &amp;amp; API sandbox, and then use &lt;strong&gt;Burp Suite&lt;/strong&gt;&amp;rsquo;s Proxy and Repeater tools to intercept, analyze, modify, and replay HTTP traffic to solve three hands-on debugging challenges.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;strong&gt;Time:&lt;/strong&gt; 45–60 minutes.
&lt;strong&gt;Difficulty:&lt;/strong&gt; ⭐⭐☆☆☆.
&lt;strong&gt;Learning Outcomes:&lt;/strong&gt; Intercept and modify outgoing client requests, bypass client-side parameters, perform ID enumeration in the Repeater, and inject custom headers.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-2/01-llm-deployment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-2/01-llm-deployment/</guid><description>&lt;h1 id="lab-deploying-private-llms-with-vllm--api-gateway"&gt;Lab: Deploying Private LLMs with vLLM &amp;amp; API Gateway&lt;a class="anchor" href="#lab-deploying-private-llms-with-vllm--api-gateway"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This practical lab guides you through serving open-source Large Language Models (LLMs) like Gemma 2B or Qwen 2.5 locally and in production using &lt;strong&gt;vLLM&lt;/strong&gt; (or Ollama/llama.cpp as a CPU fallback). To prevent data leakage and secure your infrastructure, you will build a FastAPI API Gateway that wraps your model server behind secure API keys and supports real-time OpenAI-compatible token streaming.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 Client[Web Browser / SDK] --&amp;gt;|1. Request with API Key| Gateway[FastAPI Gateway]
 Gateway --&amp;gt;|2. Validate Key| Vault[(Authorized Keys)]
 Gateway --&amp;gt;|3. Forward Request| LLMServer[vLLM Inference Engine]
 LLMServer --&amp;gt;|4. Generate Tokens| GPU[NVIDIA / AMD GPU]
 GPU --&amp;gt;|5. Compute| LLMServer
 LLMServer --&amp;gt;|6. Stream Response| Gateway
 Gateway --&amp;gt;|7. SSE Stream| Client&lt;/pre&gt;&lt;h3 id="theoretical-context-vllm--quantization"&gt;Theoretical Context: vLLM &amp;amp; Quantization&lt;a class="anchor" href="#theoretical-context-vllm--quantization"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;h4 id="why-vllm"&gt;Why vLLM?&lt;a class="anchor" href="#why-vllm"&gt;#&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Standard PyTorch or Hugging Face serving frameworks handle requests sequentially or in static batch queues. This causes massive latency and wastes GPU capacity. vLLM solves this with:&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-2/02-socket-chat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-2/02-socket-chat/</guid><description>&lt;h1 id="lab-websocket-chat-with-redis--postgresql"&gt;Lab: WebSocket Chat with Redis &amp;amp; PostgreSQL&lt;a class="anchor" href="#lab-websocket-chat-with-redis--postgresql"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This practical lab guides you through building a real-time, authenticated chat application. You will implement Google OAuth for authentication, use a PostgreSQL database to persist messages, integrate Redis to cache sessions and chat histories, and establish WebSockets for instantaneous messaging.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart TD
 Browser[Web Browser]
 FastAPI[FastAPI Backend]
 Google[Google OAuth Server]
 Postgres[(PostgreSQL Database)]
 Redis[(Redis Cache)]

 Browser --&amp;gt;|1. OAuth Redirect &amp;amp; Callback| Google
 Browser --&amp;gt;|2. HTTP Requests with Session Cookie| FastAPI
 Browser --&amp;gt;|3. Establish WebSocket Connection| FastAPI
 FastAPI --&amp;gt;|Check/Set Session &amp;amp; Cache History| Redis
 FastAPI --&amp;gt;|Query/Persist Messages &amp;amp; Users| Postgres&lt;/pre&gt;&lt;h3 id="real-developer-use-case"&gt;Real Developer Use Case&lt;a class="anchor" href="#real-developer-use-case"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;In production systems, raw WebSockets are stateful and memory-intensive. Keeping chat histories or sessions in-memory limits scalability. By storing persistent records in PostgreSQL and using Redis for high-speed session checks and chat history caches, you create a stateless, horizontally-scalable backend capable of handling thousands of concurrent users.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-3/cost-tracking-dashboard-langsmith/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-3/cost-tracking-dashboard-langsmith/</guid><description>&lt;h1 id="lab-32--cost-tracking-dashboard-via-langsmith"&gt;Lab 3.2 — Cost-Tracking Dashboard via LangSmith&lt;a class="anchor" href="#lab-32--cost-tracking-dashboard-via-langsmith"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;What you&amp;rsquo;ll build:&lt;/strong&gt; A benchmark system that runs the same task using four different prompt strategies — zero-shot, few-shot, Chain-of-Thought, and Self-Consistency — traces every call to LangSmith, then exposes a FastAPI endpoint that returns a live cost comparison dashboard.&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;┌────────────────────────────────────────────────────────────┐
│ Input: &amp;#34;What is 15% of 840?&amp;#34; │
│ │
│ Strategy 1: Zero-shot → &amp;#34;126&amp;#34; $0.000012 │
│ Strategy 2: Few-shot → &amp;#34;126&amp;#34; $0.000018 │
│ Strategy 3: CoT → &amp;#34;...= 126&amp;#34; $0.000042 │
│ Strategy 4: Self-Consist. → &amp;#34;126 (7/7)&amp;#34; $0.000294 │
│ │
│ → LangSmith shows every call with token counts &amp;amp; cost │
│ → /dashboard endpoint returns live comparison JSON │
└────────────────────────────────────────────────────────────┘&lt;/code&gt;&lt;/pre&gt;&lt;hr&gt;
&lt;h2 id="architecture"&gt;Architecture&lt;a class="anchor" href="#architecture"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;graph TD
 A[benchmark.py] --&amp;gt;|traceable| B(run_zero_shot)
 A --&amp;gt;|traceable| C(run_few_shot)
 A --&amp;gt;|traceable| D(run_cot)
 A --&amp;gt;|traceable| E(run_self_consistency)
 B --&amp;gt;|LangSmith API| F[LangSmith Cloud Dashboard]
 C --&amp;gt;|LangSmith API| F
 D --&amp;gt;|LangSmith API| F
 E --&amp;gt;|LangSmith API| F
 G[FastAPI Server] --&amp;gt;|Query API| F
 G --&amp;gt;|Endpoint: /dashboard| H[Live Cost JSON Dashboard]&lt;/pre&gt;&lt;hr&gt;
&lt;details&gt;
&lt;summary&gt;**Step 1 — Project Setup**&lt;/summary&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mkdir tds-lab-3-2
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd tds-lab-3-2
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv init --no-workspace&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Install dependencies:&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-3/youtube-subtitles-topics-json-pipeline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-3/youtube-subtitles-topics-json-pipeline/</guid><description>&lt;h1 id="lab-31--youtube--subtitles--topics--json-pipeline"&gt;Lab 3.1 — YouTube → Subtitles → Topics → JSON Pipeline&lt;a class="anchor" href="#lab-31--youtube--subtitles--topics--json-pipeline"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;What you&amp;rsquo;ll build:&lt;/strong&gt; A CLI tool that takes any YouTube URL and produces a structured JSON summary containing topics, key Q&amp;amp;A pairs, chapter breakdowns, and answer timestamps — all extracted by Claude from the auto-generated subtitles.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Input&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python pipeline.py &lt;span style="color:#e6db74"&gt;&amp;#34;https://www.youtube.com/watch?v=VIDEO_ID&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Output: output/VIDEO_ID_summary.json&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;title&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Introduction to Docker&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;duration_minutes&amp;#34;&lt;/span&gt;: 45,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;topics&amp;#34;&lt;/span&gt;: &lt;span style="color:#f92672"&gt;[&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;containerization&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;Dockerfile&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;docker-compose&amp;#34;&lt;/span&gt;, ...&lt;span style="color:#f92672"&gt;]&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;chapters&amp;#34;&lt;/span&gt;: &lt;span style="color:#f92672"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;title&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;What is Docker?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;start_time&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;00:00&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;end_time&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;05:30&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;summary&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;Docker packages apps into containers...&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;]&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;qa_pairs&amp;#34;&lt;/span&gt;: &lt;span style="color:#f92672"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;question&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;What is the difference between image and container?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;answer&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;An image is a blueprint; a container is a running instance.&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;timestamp&amp;#34;&lt;/span&gt;: &lt;span style="color:#e6db74"&gt;&amp;#34;03:42&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;hr&gt;
&lt;h2 id="architecture"&gt;Architecture&lt;a class="anchor" href="#architecture"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;graph TD
 A[YouTube URL] --&amp;gt;|yt-dlp| B(VTT Subtitle File)
 B --&amp;gt;|vtt_parser.py| C(Timestamped Segments)
 C --&amp;gt;|Claude Haiku| D[Topic Extraction]
 C --&amp;gt;|Claude Haiku| E[Chapter Segmentation]
 C --&amp;gt;|Claude Haiku| F[Q&amp;amp;A Extraction]
 D --&amp;gt; G[Assemble VideoSummary JSON]
 E --&amp;gt; G
 F --&amp;gt; G&lt;/pre&gt;&lt;hr&gt;
&lt;details&gt;
&lt;summary&gt;**Step 1 — Project Setup**&lt;/summary&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mkdir tds-lab-3-1
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd tds-lab-3-1
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv init --no-workspace&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Install dependencies:&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-4/capstone-bs-degree-chatbot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-4/capstone-bs-degree-chatbot/</guid><description>&lt;h1 id="capstone--bs-degree-chatbot"&gt;Capstone — BS Degree Chatbot&lt;a class="anchor" href="#capstone--bs-degree-chatbot"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="objective"&gt;Objective&lt;a class="anchor" href="#objective"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Build a production-grade RAG chatbot that answers user questions about the IIT Madras BS Data Science programme using official documentation.&lt;/p&gt;
&lt;h2 id="requirements"&gt;Requirements&lt;a class="anchor" href="#requirements"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Scrape academics, courses, and qualifier pages from the official IIT Madras BS program site.&lt;/li&gt;
&lt;li&gt;Build an ingestion pipeline using a vector database (e.g., Qdrant) with contextual chunking.&lt;/li&gt;
&lt;li&gt;Implement a hybrid retrieval strategy combining dense vector embeddings and BM25 sparse index, merged via RRF and an optional reranker.&lt;/li&gt;
&lt;li&gt;Build a FastAPI backend with a &lt;code&gt;/chat&lt;/code&gt; endpoint and a Streamlit-based UI.&lt;/li&gt;
&lt;li&gt;Setup a RAGAS evaluation suite to measure retrieval and generation accuracy across different retrieval strategies.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="deliverables"&gt;Deliverables&lt;a class="anchor" href="#deliverables"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Source code.&lt;/li&gt;
&lt;li&gt;Link to the live Streamlit chatbot application.&lt;/li&gt;
&lt;li&gt;RAGAS evaluation report comparing Naive, Hybrid, and Contextual RAG.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/labs/week-4/ragas-evaluation-dashboard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-4/ragas-evaluation-dashboard/</guid><description>&lt;h1 id="lab--ragas-evaluation-dashboard"&gt;Lab — RAGAS Evaluation Dashboard&lt;a class="anchor" href="#lab--ragas-evaluation-dashboard"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="objective"&gt;Objective&lt;a class="anchor" href="#objective"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Build an interactive Streamlit dashboard that runs RAGAS evaluation across multiple RAG retrieval strategies side-by-side.&lt;/p&gt;
&lt;h2 id="requirements"&gt;Requirements&lt;a class="anchor" href="#requirements"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Implement three RAG strategies: Naive RAG (dense only), Hybrid RAG (dense + BM25 + RRF), and Contextual RAG (using Anthropic&amp;rsquo;s chunk context injection).&lt;/li&gt;
&lt;li&gt;Generate a set of test questions and ground-truth answers using an LLM.&lt;/li&gt;
&lt;li&gt;Evaluate the retrieval strategies using RAGAS metrics (Faithfulness, Context Recall, Factual Correctness, Response Relevancy).&lt;/li&gt;
&lt;li&gt;Build a Streamlit dashboard showing metrics comparison tables, strategy charts, and per-question breakdowns.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="deliverables"&gt;Deliverables&lt;a class="anchor" href="#deliverables"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Source code.&lt;/li&gt;
&lt;li&gt;Screenshot of the Streamlit dashboard comparing all three strategies.&lt;/li&gt;
&lt;li&gt;Downloadable evaluation results JSON file.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/labs/week-5/capstone-autonomous-research-agent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-5/capstone-autonomous-research-agent/</guid><description>&lt;h1 id="capstone--autonomous-daily-ai-research-blog"&gt;Capstone — Autonomous Daily AI Research Blog&lt;a class="anchor" href="#capstone--autonomous-daily-ai-research-blog"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Build a static blog where a LangChain agent researches recent AI news every morning, writes a short cited digest, remembers previous coverage, commits the new Markdown post, and deploys the updated site to GitHub Pages.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;This is a &lt;strong&gt;single-agent loop&lt;/strong&gt;, not a multi-agent demo. The agent can only search; normal Python code validates and writes the post. This keeps the project small, understandable, and safer.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-6/capstone-ai-signature-detection-cropper/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-6/capstone-ai-signature-detection-cropper/</guid><description>&lt;h1 id="capstone--ai-signature-detection--cropper"&gt;Capstone — AI Signature Detection &amp;amp; Cropper&lt;a class="anchor" href="#capstone--ai-signature-detection--cropper"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Find handwritten signatures on scanned documents, crop them out, verify them with a second model, and be honest about your error rate.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~6–8 hours
🔗 needs: &lt;a href="/2026-05/week-6/vision-models-for-scraping/"&gt;Vision Models for Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/image-processing-pipeline/"&gt;Image Processing Pipeline&lt;/a&gt; · &lt;a href="/2026-05/week-6/document-parsing/"&gt;Document Parsing&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A realistic document-AI pipeline: detect a region, crop it, and use a &lt;strong&gt;second, independent&lt;/strong&gt; model to check the first one. The interesting engineering is in the verification and the measurement, not the detection call.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-6/capstone-job-posting-scraper-tracker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-6/capstone-job-posting-scraper-tracker/</guid><description>&lt;h1 id="capstone--job-posting-scraper--tracker"&gt;Capstone — Job Posting Scraper &amp;amp; Tracker&lt;a class="anchor" href="#capstone--job-posting-scraper--tracker"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Build a scraper that runs itself, notices what changed, and answers questions about the market.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~6–8 hours
🔗 needs: &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-6/change-detection-dedup/"&gt;Change Detection &amp;amp; Dedup&lt;/a&gt; · &lt;a href="/2026-05/week-6/duckdb-parquet/"&gt;DuckDB + Parquet&lt;/a&gt; · &lt;a href="/2026-05/week-6/scheduled-scraping/"&gt;Scheduled Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Anyone can scrape a page once. This capstone is about the harder, more valuable thing: a pipeline that runs unattended for weeks, doesn&amp;rsquo;t duplicate data, doesn&amp;rsquo;t get banned, and produces a dataset worth querying.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-6/capstone-live-multilingual-travel-translator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-6/capstone-live-multilingual-travel-translator/</guid><description>&lt;h1 id="capstone--live-multilingual-travel-translator"&gt;Capstone — Live Multilingual Travel Translator&lt;a class="anchor" href="#capstone--live-multilingual-travel-translator"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Speech in, speech out, in another language — a three-model pipeline where every stage can fail, and your job is to make the whole thing not fall over.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~6–8 hours
🔗 needs: &lt;a href="/2026-05/week-6/speech-ai/"&gt;Speech AI&lt;/a&gt; · &lt;a href="/2026-05/week-2/01-fastapi/"&gt;FastAPI&lt;/a&gt; · &lt;a href="/2026-05/week-3/structured-output/"&gt;Structured Output&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Chaining STT → translation → TTS is easy to demo and hard to make reliable: errors compound, latency stacks, and each stage has its own failure mode. This capstone is graded on the engineering around the models.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-6/osint-open-source-dossier/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-6/osint-open-source-dossier/</guid><description>&lt;h1 id="lab--open-source-organisation-dossier"&gt;Lab — Open-Source Organisation Dossier&lt;a class="anchor" href="#lab--open-source-organisation-dossier"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Assemble a fully-cited profile of an organisation&amp;rsquo;s public footprint, where every single claim carries its source, timestamp, and confidence level.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~3–4 hours
🔗 needs: &lt;a href="/2026-05/week-6/osint/"&gt;OSINT — Infrastructure &amp;amp; Records&lt;/a&gt; · &lt;a href="/2026-05/week-6/google-dork/"&gt;Google Dorking&lt;/a&gt; · &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The skill being graded is &lt;strong&gt;not&lt;/strong&gt; collection — anyone can paste tool output. It&amp;rsquo;s &lt;strong&gt;verification&lt;/strong&gt;: knowing what your evidence actually supports, and saying so honestly.&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;⚖️ &lt;strong&gt;Scope rules — a submission that breaks any of these scores zero.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-6/scheduled-scraper-github-actions-cron/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-6/scheduled-scraper-github-actions-cron/</guid><description>&lt;h1 id="lab--scheduled-scraper-with-github-actions"&gt;Lab — Scheduled Scraper with GitHub Actions&lt;a class="anchor" href="#lab--scheduled-scraper-with-github-actions"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Get a scraper running on a free daily cron that commits its own data — the foundation every other Week 6 lab builds on.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~2 hours
🔗 needs: &lt;a href="/2026-05/week-6/scheduled-scraping/"&gt;Scheduled Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/duckdb-parquet/"&gt;DuckDB + Parquet&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A small, complete pipeline: fetch → store → schedule → query. Deliberately uses a &lt;strong&gt;public, documented API&lt;/strong&gt; so nothing here is ethically ambiguous.&lt;/p&gt;
&lt;h2 id="objective"&gt;Objective&lt;a class="anchor" href="#objective"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Collect Hacker News top stories daily, accumulate them as Parquet, and query the result with DuckDB.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-7/01-red-team-your-api-guardrails/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-7/01-red-team-your-api-guardrails/</guid><description>&lt;h1 id="lab--red-team-your-own-llm-api"&gt;Lab — Red-Team Your Own LLM API&lt;a class="anchor" href="#lab--red-team-your-own-llm-api"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Attack a system you built, prove which defence stopped each attack, and leave behind a regression suite that keeps it fixed.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~4–5 hours
🔗 needs: &lt;a href="/2026-05/week-7/03-llm-security-offensive/"&gt;LLM Security — Offensive&lt;/a&gt; · &lt;a href="/2026-05/week-7/04-llm-safety-defensive/"&gt;LLM Safety — Defensive&lt;/a&gt; · &lt;a href="/2026-05/week-7/05-owasp-llm-top-10/"&gt;OWASP LLM Top 10&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll build a deliberately weak LLM service, break it, harden it, and then prove the hardening works — with tests that run in CI.&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;⚖️ &lt;strong&gt;Target only your own deployment.&lt;/strong&gt; Every attack in this lab runs against the service &lt;em&gt;you&lt;/em&gt; build and host. Probing a third party&amp;rsquo;s LLM product is unauthorised testing. If you want a harder target, make your own service harder.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-7/02-full-cicd-cloud-run/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-7/02-full-cicd-cloud-run/</guid><description>&lt;h1 id="lab--full-cicd-pipeline-to-cloud-run"&gt;Lab — Full CI/CD Pipeline to Cloud Run&lt;a class="anchor" href="#lab--full-cicd-pipeline-to-cloud-run"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;One &lt;code&gt;git push&lt;/code&gt; → tests, image build, deploy, health check. Plus the parts people skip: a rollback, a cost cap, and an approval gate.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~4–5 hours
🔗 needs: &lt;a href="/2026-05/week-7/01-github-actions-advanced/"&gt;GitHub Actions Advanced&lt;/a&gt; · &lt;a href="/2026-05/week-7/02-advanced-docker/"&gt;Advanced Docker&lt;/a&gt; · &lt;a href="/2026-05/week-7/07-serverless-functions/"&gt;Serverless Functions&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Everything in Week 7 assembled into one working pipeline.&lt;/p&gt;
&lt;h2 id="what-youre-building"&gt;What you&amp;rsquo;re building&lt;a class="anchor" href="#what-youre-building"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 P[&amp;#34;git push&amp;#34;] --&amp;gt; T[&amp;#34;Test matrix + lint&amp;#34;]
 T --&amp;gt; B[&amp;#34;Build image&amp;lt;br/&amp;gt;(cached, multi-stage)&amp;#34;]
 B --&amp;gt; SC[&amp;#34;Scan for CVEs + secrets&amp;#34;]
 SC --&amp;gt; ST[&amp;#34;Deploy to STAGING&amp;#34;]
 ST --&amp;gt; H[&amp;#34;Smoke test /health&amp;#34;]
 H --&amp;gt;|&amp;#34;main only&amp;#34;| A{&amp;#34;Manual approval&amp;#34;}
 A --&amp;gt; PR[&amp;#34;Deploy to PRODUCTION&amp;#34;]
 PR --&amp;gt; V[&amp;#34;Verify + auto-rollback on failure&amp;#34;]&lt;/pre&gt;&lt;h2 id="requirements"&gt;Requirements&lt;a class="anchor" href="#requirements"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;1. The app.&lt;/strong&gt; Any small FastAPI service with &lt;code&gt;/health&lt;/code&gt; and one real endpoint. Reuse your Week 6 scraper API or the Week 7 hardened LLM service.&lt;/p&gt;</description></item><item><title/><link>/2026-05/labs/week-8/09-full-gcp-walkthrough/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/labs/week-8/09-full-gcp-walkthrough/</guid><description>&lt;h1 id="milestone-full-gcp-walkthrough"&gt;Milestone: Full GCP Walkthrough&lt;a class="anchor" href="#milestone-full-gcp-walkthrough"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;!-- Week 8 milestone walkthrough. --&gt;
&lt;p&gt;&lt;a href="https://youtu.be/lvZk_sc8u5I?si=EBox9sm7ygXswiwL"&gt;&lt;img src="https://img.youtube.com/vi/lvZk_sc8u5I/0.jpg" alt="GCP walkthrough" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;GCP is not one big computer in the cloud. It is a set of managed services, identities, regions, bills, and logs. Your job is to connect only the few pieces your product needs—and be able to explain every connection.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~18 min read · ~45 min guided walkthrough
🔗 needs: &lt;a href="/2026-05/week-8/01-cloud-storage-ml/"&gt;Cloud Storage for ML&lt;/a&gt; · &lt;a href="/2026-05/week-8/02-bigquery-ml/"&gt;BigQuery ML&lt;/a&gt; · &lt;a href="/2026-05/week-7/09-cost-alerting/"&gt;Cost Alerting &amp;amp; Budgets&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/large-language-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/large-language-models/</guid><description>&lt;h1 id="large-language-models"&gt;Large Language Models&lt;a class="anchor" href="#large-language-models"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This module covers the practical usage of large language models (LLMs).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;LLMs incur a cost.&lt;/strong&gt; For the May 2025 batch, use &lt;a href="https://aipipe.org/"&gt;aipipe.org&lt;/a&gt; as a proxy.
Emails with &lt;code&gt;@ds.study.iitm.ac.in&lt;/code&gt; get a &lt;strong&gt;$1 per calendar month&lt;/strong&gt; allowance. (Don&amp;rsquo;t exceed that.)&lt;/p&gt;
&lt;p&gt;Read the &lt;a href="https://github.com/sanand0/aipipe"&gt;AI Pipe documentation&lt;/a&gt; to learn how to use it. But in short:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Replace &lt;code&gt;OPENAI_BASE_URL&lt;/code&gt;, i.e. &lt;code&gt;https://api.openai.com/v1&lt;/code&gt; with &lt;code&gt;https://aipipe.org/openrouter/v1...&lt;/code&gt; or &lt;code&gt;https://aipipe.org/openai/v1...&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Replace &lt;code&gt;OPENAI_API_KEY&lt;/code&gt; with the &lt;a href="https://aipipe.org/login"&gt;&lt;code&gt;AIPIPE_TOKEN&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Replace model names, e.g. &lt;code&gt;gpt-4.1-nano&lt;/code&gt;, with &lt;code&gt;openai/gpt-4.1-nano&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For example, let&amp;rsquo;s use &lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-0-flash-lite"&gt;Gemini 2.0 Flash Lite&lt;/a&gt; via &lt;a href="https://openrouter.ai/google/gemini-2.0-flash-lite-001"&gt;OpenRouter&lt;/a&gt; for chat completions and &lt;a href="https://platform.openai.com/docs/models/text-embedding-3-small"&gt;Text Embedding 3 Small&lt;/a&gt; via &lt;a href="https://platform.openai.com/docs/"&gt;OpenAI&lt;/a&gt; for embeddings:&lt;/p&gt;</description></item><item><title/><link>/2026-05/live-sessions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/live-sessions/</guid><description>&lt;h1 id="live-sessions"&gt;Live Sessions&lt;a class="anchor" href="#live-sessions"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Live sessions by the instructors and TAs are recorded and uploaded to the &lt;a href="https://www.youtube.com/@se-lr5ff/videos"&gt;YouTube Tools in Data Science channel&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;These are transcribed by &lt;a href="https://github.com/sanand0/tools-in-data-science-public/blob/main/live-sessions/faq.sh"&gt;faq.sh&lt;/a&gt; into FAQ-style markdown files, which are stored here.&lt;/p&gt;
&lt;!-- The section below is generated by faq.sh and should not be edited manually. --&gt;
&lt;!--

- [x] Header
- [ ] Duration
- [ ] Cookies.txt if it exists -- cookie editor extension: https://microsoftedge.microsoft.com/addons/detail/cookiemanager-cookie-ed/mmegchnodbbdfhhccbnnbalnedndcbil?utm_source=chatgpt.com
- [ ] Prettier
 --&gt;
&lt;h2 id="sessions"&gt;Sessions&lt;a class="anchor" href="#sessions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://youtu.be/UlUf6roSG3c"&gt;Video&lt;/a&gt; | 2025 10 04 Project 1 - Session 1 TDS Sep 2025 (2h 24m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/0x5WK-Haz7Y"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20251004-0x5wk-haz7y/"&gt;FAQ&lt;/a&gt; | 2025 10 03 week 2 - Session 4 TDS Sep 2025 (2h 16m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/dZtLGbQ_EW8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20251004-dztlgbq_ew8/"&gt;FAQ&lt;/a&gt; | 2025 09 30 week 2 - Session 3 TDS Sep 2025 (1h 45m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/4ojO3e2PmP8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250929-4ojo3e2pmp8/"&gt;FAQ&lt;/a&gt; | 2025 09 28 week 2 - Session 2 TDS Sep 2025 (2h 36m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ytGxUwBkcXU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250929-ytgxuwbkcxu/"&gt;FAQ&lt;/a&gt; | 2025 09 27 week 2 - Session 1 TDS Sep 2025 (2h 53m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/vBqJUhbZU3c"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250929-vbqjuhbzu3c/"&gt;FAQ&lt;/a&gt; | 2025 09 24 week 1 - Session 5 TDS Sep 2025 (1h 41m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/thACpPPTR1I"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250929-thacppptr1i/"&gt;FAQ&lt;/a&gt; | 2025 09 29 Week 1 - Session 4 TDS Sep 2025 (1h 38m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/RoRy1jKewhA"&gt;Video&lt;/a&gt; | 2025 09 21 Week 1 - Session 2 TDS Sep 2025 (2h 37m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/sXlsRhw5X94"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250922-sxlsrhw5x94/"&gt;FAQ&lt;/a&gt; | 2025 09 22 Week 1 - Session 3 TDS Sep 2025 (1h 58m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/B7iHBHCruCc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250921-b7ihbhcrucc/"&gt;FAQ&lt;/a&gt; | 2025 09 20 Week 1 - Session 1 TDS Sep 2025 (2h 52m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/FXZzXm2_3KI"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250806-fxzzxm2_3ki/"&gt;FAQ&lt;/a&gt; | 2025 08 04 project 2 - Q&amp;amp;A 2 TDS May 2025 (1h 59m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/PZMAApzlJCo"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250729-pzmaapzljco/"&gt;FAQ&lt;/a&gt; | 2025 07 26 project 2 - Session 2 TDS May 2025 (1h 42m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/p_OgYP-6IdU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250725-p_ogyp-6idu/"&gt;FAQ&lt;/a&gt; | 2025 07 24 project 2 - Session 1 TDS May 2025 (1h 50m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/miNphxMz9wc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250722-minphxmz9wc/"&gt;FAQ&lt;/a&gt; | 2025 07 21 project 2 - Q&amp;amp;A 1 TDS May 2025 (1h 32m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/afaqw038F4Q"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250719-afaqw038f4q/"&gt;FAQ&lt;/a&gt; | 2025 07 18 MOCK ROE session 4 - TDS May 2025 (2h 7m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/bS-ZFSTHoqI"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250719-bs-zfsthoqi/"&gt;FAQ&lt;/a&gt; | 2025 07 17 MOCK ROE session 3 - TDS May 2025 (1h 33m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/4ZM-krZjvDQ"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250719-4zm-krzjvdq/"&gt;FAQ&lt;/a&gt; | 2025 07 16 MOCK ROE session 2 - TDS May 2025 (0h 53m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/oHDT0CjHpr8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250711-ohdt0cjhpr8/"&gt;FAQ&lt;/a&gt; | 2025 07 10 MOCK ROE session1 - TDS May 2025 (1h 56m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ijia1cLUQck"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250707-ijia1cluqck/"&gt;FAQ&lt;/a&gt; | 2025 07 06 Week 6 Session 2 - TDS May 2025 (1h 49m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/kY27w7bXyac"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250706-ky27w7bxyac/"&gt;FAQ&lt;/a&gt; | 2025 07 05 Week 6 Session 1 - TDS May 2025 (1h 17m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/eFa-LofjwfM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250629-efa-lofjwfm/"&gt;FAQ&lt;/a&gt; | 2025 06 28 Week 5 Session 1 - TDS May 2025 (2h 8m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/RFIsZSZhqGs"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250621-rfiszszhqgs/"&gt;FAQ&lt;/a&gt; | 2025 06 15 Week 4 1 Session 1 - TDS May 2025 (1h 26m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/VZ05EVdj_Dk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250621-vz05evdj_dk/"&gt;FAQ&lt;/a&gt; | 2025 06 12 Project 1 - DCS TDS May 2025 (1h 2m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/q8cZ2b7WCHI"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250612-q8cz2b7wchi/"&gt;FAQ&lt;/a&gt; | 2025 06 12 Project 1 Session 3 - TDS May 2025 (2h 26m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/PwYU_vRmcl0"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250612-pwyu_vrmcl0/"&gt;FAQ&lt;/a&gt; | 2025 06 11 Project 1 Session 2 - TDS May 2025 (2h 59m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/eeLLWOVSc8M"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250611-eellwovsc8m/"&gt;FAQ&lt;/a&gt; | 2025 06 10 Project 1 Session 1 - TDS May 2025 (2h 28m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/PYj9eOPpH7k"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250531-pyj9eopph7k/"&gt;FAQ&lt;/a&gt; | 2025 05 29 Week 3 Session 4 - TDS May 2025 (2h 44m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/qqgBvapfrCc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250531-qqgbvapfrcc/"&gt;FAQ&lt;/a&gt; | 2025 05 28 Week 3 Session 3 - TDS May 2025 (1h 30m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/oWAaWzmOwmM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250531-owaawzmowmm/"&gt;FAQ&lt;/a&gt; | 2025 05 25 Week 3 Session 2 - TDS May 2025 (1h 38m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/s-ZQ2I1lE7Y"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250524-s-zq2i1le7y/"&gt;FAQ&lt;/a&gt; | 2025 05 24 Week 3 Session 1 - TDS May 2025 (1h 47m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/JXiJVWisKns"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250524-jxijvwiskns/"&gt;FAQ&lt;/a&gt; | 2025 05 23 Week 2 Session 2 - Github Actions, Docker - TDS May 2025 (3h 17m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/iMLGnriiecc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250522-imlgnriiecc/"&gt;FAQ&lt;/a&gt; | 2025 05 22 Week 2 Session 2 - Local Lamma, Ngrok, Google Authentication TDS May 2025 (2h 40m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Ptq0Me5tmFk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250522-ptq0me5tmfk/"&gt;FAQ&lt;/a&gt; | 2025 05 21 Week 2 Session 1 TDS May 2025 (2h 31m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/1a2x_oDwHlk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250517-1a2x_odwhlk/"&gt;FAQ&lt;/a&gt; | 2025 05 17 Week 1 Session 5 TDS May 2025 (2h 37m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/c0GXO3dh_sM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250517-c0gxo3dh_sm/"&gt;FAQ&lt;/a&gt; | 2025 05 16 Week 1 Session 4 Env Settings, Env Variables, git and github TDS May 2025 (2h 12m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/BCF4oFqE8Tw"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250517-bcf4ofqe8tw/"&gt;FAQ&lt;/a&gt; | 2025 05 15 Week 1 Session 3 Linux, BASH, git &amp;amp; github TDS May 2025 (3h 3m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/TiqtXjmmKFA"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250515-tiqtxjmmkfa/"&gt;FAQ&lt;/a&gt; | 2025 05 11 Week 1 Session 2 Setup &amp;amp; Debugging TDS May 2025 (3h 3m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/iuYPkqjfvM8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250512-iuypkqjfvm8/"&gt;FAQ&lt;/a&gt; | 2025-05-10 Week 1 - Session 1: Introduction to TDS - TDS May 25 (2h 29m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Er49fGwYnPc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250410-er49fgwynpc/"&gt;FAQ&lt;/a&gt; | 2025-04-09 Live Session - Qs &amp;amp; As - TDS Jan 25 (1h 19m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/CXRnZBaR4BY"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250409-cxrnzbar4by/"&gt;FAQ&lt;/a&gt; | 2025-04-08 End Term Session 1 - TDS Jan 25 (1h 59m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Nrm9gor3Thk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250329-nrm9gor3thk/"&gt;FAQ&lt;/a&gt; | 2025-03-27 Week 11 - Session 3 - TDS Jan 25 (2h 7m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Kwvqd1k-VUU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250327-kwvqd1k-vuu/"&gt;FAQ&lt;/a&gt; | 2025-03-26 Week 11 - Session 2 - TDS Jan 25 (1h 27m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/WyO9CYWiP20"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250327-wyo9cywip20/"&gt;FAQ&lt;/a&gt; | 2025-03-25 Week 11 - Session 1 - TDS Jan 25 (1h 17m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/l-F_A8P-XSs"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250324-l-f_a8p-xss/"&gt;FAQ&lt;/a&gt; | 2025-03-20 Week 10 - Session 3 - TDS Jan 25 (1h 44m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/z3rta0JBeBc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250320-z3rta0jbebc/"&gt;FAQ&lt;/a&gt; | 2025-03-18 Week 10 - Session 1 - TDS Jan 25 (2h 24m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/cRg4VDBFr2k"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250314-crg4vdbfr2k/"&gt;FAQ&lt;/a&gt; | 2025-03-13 Week 9 - Session 3 - TDS Jan 25 (1h 58m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/cLJQej-Mq-Y"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250314-cljqej-mq-y/"&gt;FAQ&lt;/a&gt; | 2025-03-11 Week 9 - Session 1 - TDS Jan 25 (2h 11m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/WG_eZaACctM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250314-wg_ezaacctm/"&gt;FAQ&lt;/a&gt; | 2025-03-12 Week 9 - Session 2 - TDS Jan 25 (1h 14m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/zc8PDm8lMhE"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250309-zc8pdm8lmhe/"&gt;FAQ&lt;/a&gt; | 2025-03-04 Week 8 - Session 1 - TDS Jan 25 (1h 52m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/mJM-IfeLq1Q"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250309-mjm-ifelq1q/"&gt;FAQ&lt;/a&gt; | 2025-03-04 Week 8 - Session 3 - TDS Jan 25 (1h 52m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/4P1204s1T1o"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250309-4p1204s1t1o/"&gt;FAQ&lt;/a&gt; | 2025-03-04 Week 8 - Session 2 - TDS Jan 25 (0h 47m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/oL-cqokYE2A"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250228-ol-cqokye2a/"&gt;FAQ&lt;/a&gt; | 2025-02-27 Week 7 - Session 3 - TDS Jan 25 (2h 24m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/nrk19yJnXgM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250228-nrk19yjnxgm/"&gt;FAQ&lt;/a&gt; | 2025-02-26 Week 7 - Session 2 - TDS Jan 25 (5h 8m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/l2zUly456eM"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250221-l2zuly456em/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 11 10 19 49 (1h 39m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/6aOAB9L7Xjg"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250221-6aoab9l7xjg/"&gt;FAQ&lt;/a&gt; | 2025-02-18 Week 6 - Session 3 - TDS Jan 25 (1h 34m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/1UbY1R4lrFk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250219-1uby1r4lrfk/"&gt;FAQ&lt;/a&gt; | 2025-02-18 Week 6 - Session 1 - TDS Jan 25 (2h 8m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/NMnwKp5tR-w"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250214-nmnwkp5tr-w/"&gt;FAQ&lt;/a&gt; | 2025-02-13 Week 5 - Session 3 - TDS Jan 25 (3h 12m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/jXj6bqy4R4c"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250212-jxj6bqy4r4c/"&gt;FAQ&lt;/a&gt; | 2025-02-11 Week 5 - Session 1 - TDS Jan 25 (3h 57m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/SiW-rcMk0Nk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250207-siw-rcmk0nk/"&gt;FAQ&lt;/a&gt; | 2025-02-06 Week 4 - Session 4 - TDS Jan 25 (2h 16m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/u5RFmePd7NQ"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250206-u5rfmepd7nq/"&gt;FAQ&lt;/a&gt; | 2025-02-06 Week 4 - Session 3 - TDS Jan 25 (2h 24m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/8A7Z_PN_PzQ"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250204-8a7z_pn_pzq/"&gt;FAQ&lt;/a&gt; | 2025-02-04 Week 4 - Session 1 - TDS Jan 25 (1h 27m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/tsn7B7mDzw8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250201-tsn7b7mdzw8/"&gt;FAQ&lt;/a&gt; | 2025-02-01 Week 3 - Session 5 - TDS Jan 25 (2h 34m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/sdg4N-H4BR0"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250131-sdg4n-h4br0/"&gt;FAQ&lt;/a&gt; | 2025-01-31 Week 3 - Session 4 - TDS Jan 25 (2h 45m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/6VfrL5b8lLc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250130-6vfrl5b8llc/"&gt;FAQ&lt;/a&gt; | 2025-01-30 Week 3 - Session 3 - TDS Jan 25 (2h 35m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/EPiVIP97fzI"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250129-epivip97fzi/"&gt;FAQ&lt;/a&gt; | 2025-01-29 Week 3 - Session 2 - TDS Jan 25 (1h 28m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/lmSMQ5LWa30"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250128-lmsmq5lwa30/"&gt;FAQ&lt;/a&gt; | 2025-01-28 Week 3 - Session 1 - TDS Jan 25 (2h 0m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/TxGY540ru3A"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250123-txgy540ru3a/"&gt;FAQ&lt;/a&gt; | 2025-01-23 Week 2 - Session 4 - TDS Jan 25 (2h 5m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/QnLi-C_LiXk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250122-qnli-c_lixk/"&gt;FAQ&lt;/a&gt; | 2025-01-22 Week 2 - Session 3 - TDS Jan 25 (1h 19m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/0e0RhXREnxU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250121-0e0rhxrenxu/"&gt;FAQ&lt;/a&gt; | 2025-01-21 Week 2 - Session 2 - TDS Jan 25 (2h 15m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/aJnygTpma7M"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250120-ajnygtpma7m/"&gt;FAQ&lt;/a&gt; | 2025-01-20 Week 2 - Session 1 - TDS Jan 25 (2h 3m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/hG5WqtbpfkI"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250117-hg5wqtbpfki/"&gt;FAQ&lt;/a&gt; | 2025-01-17 Week 1 - Session 2 - TDS Jan 25 (2h 17m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/1H5Aq7HjqwQ"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250116-1h5aq7hjqwq/"&gt;FAQ&lt;/a&gt; | 2025-01-16 Week 1 - Session 1 - TDS Jan 25 (1h 59m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/VTBwpPT3A3U"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20250115-vtbwppt3a3u/"&gt;FAQ&lt;/a&gt; | 2025-01-15 Introduction to Tools for Data Science Jan 25 (1h 18m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/HsfxebvUAH4"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241220-hsfxebvuah4/"&gt;FAQ&lt;/a&gt; | Mock QP - TA Session - TDS 2024 12 19 18 51 (2h 25m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ob_OO9chM-o"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241211-ob_oo9chm-o/"&gt;FAQ&lt;/a&gt; | Project 2 TA&amp;rsquo;s Session TDS 2024 12 05 19 52 (3h 0m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/gJyrKw3ThXo"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241211-gjyrkw3thxo/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS 2024 12 05 19 52 GMT+05 30 Recording (3h 0m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/TPh1srfrWDc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241117-tph1srfrwdc/"&gt;FAQ&lt;/a&gt; | ROE Session TDS 2024 11 16 19 52 (1h 13m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ifr24jckGxU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241116-ifr24jckgxu/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 11 07 19 58 GMT+05 30 – How to hack Project 1 and possibly ROE, pdf scrapi (2h 14m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/FOCsSa7nRCg"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241116-focssa7nrcg/"&gt;FAQ&lt;/a&gt; | ROE Session TDS 2024 11 15 19 53 (2h 0m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/hx0BpZT1ILQ"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241115-hx0bpzt1ilq/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 11 10 19 49 GMT+05 30 – Recording (1h 39m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/FU3cuppNuy8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241115-fu3cuppnuy8/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS 2024 11 14 19 54 (1h 52m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/7tggFDMXWi8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241105-7tggfdmxwi8/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 11 04 16 56 IST – Recording (1h 55m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/V97UfXmtTOU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241031-v97ufxmttou/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 10 27 19 55 IST – Recording (1h 45m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/LaR6fK22HEw"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241027-lar6fk22hew/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 10 24 19 50 IST – Recording 2 (2h 41m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/oqXfJKa8_ro"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241022-oqxfjka8_ro/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024-10-20 19:50 IST – Recording (1h 59m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/lV3i2fB8wbk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241022-lv3i2fb8wbk/"&gt;FAQ&lt;/a&gt; | TAs Session - 2024 10 17 20:09:44 (1h 50m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Z31q3IYTTps"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241017-z31q3iyttps/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 10 13 19 49 IST – Recording (2h 11m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Ery2oia9vMo"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241016-ery2oia9vmo/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 10 06 19 54 IST – Recording (1h 35m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/PNBZa74Gx7s"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20241002-pnbza74gx7s/"&gt;FAQ&lt;/a&gt; | TA&amp;rsquo;s Session TDS – 2024 09 29 19 49 IST – Recording (1h 40m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/NbrQTc03k04"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240830-nbrqtc03k04/"&gt;FAQ&lt;/a&gt; | TA Session 29th August: Revision: Modules 7 to 9, CSS selectors, Javascript (1h 25m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/EXo2DTLgxfU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240826-exo2dtlgxfu/"&gt;FAQ&lt;/a&gt; | TA Session 25th August: Revision: Modules 4 to 6 (1h 51m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/-_6312YVVM8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240824--_6312yvvm8/"&gt;FAQ&lt;/a&gt; | TA Session 22 August: Mock and End Term, Revision: Modules 1 to 3 (2h 6m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Gf85Cw2hwEU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240820-gf85cw2hweu/"&gt;FAQ&lt;/a&gt; | TA Session 19 August: Streamlit, Heroku (1h 47m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/fpxvSk7Qa_I"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240720-fpxvsk7qa_i/"&gt;FAQ&lt;/a&gt; | TA Session 19 July: ROE Survival : Part 2 (2h 19m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/23Ty7oaEtq4"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240719-23ty7oaetq4/"&gt;FAQ&lt;/a&gt; | TA Session 18 July: The best ROE Survival Strategies, Tricks and Tips: Part 1 (2h 12m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ELZf0n_0u9w"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240715-elzf0n_0u9w/"&gt;FAQ&lt;/a&gt; | &amp;ldquo;TA session 12 July : week 5 , LLM API and embedding&amp;rdquo; (1h 54m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/MQgOy5RNNz0"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240715-mqgoy5rnnz0/"&gt;FAQ&lt;/a&gt; | TA session 14 July : ROE intro and strategy (1h 15m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/3OdReZsvi2w"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240712-3odrezsvi2w/"&gt;FAQ&lt;/a&gt; | TA Session 11 July: LLM APi Setup and Use (2h 48m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/ce18AYM71Y8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240708-ce18aym71y8/"&gt;FAQ&lt;/a&gt; | TA Session 7 July: Project Discussion, and Excel Filters, Slope Regression (1h 51m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/97BfmXO0i7E"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240705-97bfmxo0i7e/"&gt;FAQ&lt;/a&gt; | TA Session 4 July: More Pandas, Parquet processing, Forecast.ETS, Database Processing and Week 4 GA (1h 48m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/Mmj6GE-zymE"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240701-mmj6ge-zyme/"&gt;FAQ&lt;/a&gt; | TA Session 30 June : More complex webscraping (2h 5m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/D0b4n1k2K5I"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240701-d0b4n1k2k5i/"&gt;FAQ&lt;/a&gt; | TA Session 27 June : Processing files with Pandas, Open Refine and Week 3 Graded Assignment (1h 51m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/uKvonP6TKzk"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240625-ukvonp6tkzk/"&gt;FAQ&lt;/a&gt; | TA session 20 June : Week 2 concepts ,API, JSON, PDF scraping and GA discussion (2h 32m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/KkHpfaIJN90"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240625-kkhpfaijn90/"&gt;FAQ&lt;/a&gt; | A session 23 June : Web scraping with cookies and understanding JSON (1h 57m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/IZe16QtSYtw"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240618-ize16qtsytw/"&gt;FAQ&lt;/a&gt; | TA Session 16 June : Introduction to HTML and web scraping (2h 7m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/8DC2d5sLgno"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240614-8dc2d5slgno/"&gt;FAQ&lt;/a&gt; | TA session 13 June : Introduction and week 1,2 recap (2h 2m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/356zqogHPnw"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240426-356zqoghpnw/"&gt;FAQ&lt;/a&gt; | ET revision : Attribute related to Anonymization and autocorrelation plot (1h 2m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/BSFnjDbCIJE"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240421-bsfnjdbcije/"&gt;FAQ&lt;/a&gt; | TA session : PYQ April 23 discussion part - 2 (2h 9m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/TxCA22mFuXc"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240420-txca22mfuxc/"&gt;FAQ&lt;/a&gt; | TA session : PYQ April 23 discussion part - 1 (1h 44m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/JdxqIUBmuic"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240419-jdxqiubmuic/"&gt;FAQ&lt;/a&gt; | TA session : PYQ Dec 23 discussion (1h 12m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/2miaPi6k2w0"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240308-2miapi6k2w0/"&gt;FAQ&lt;/a&gt; | 05 ROE revision session jan 24 (2h 56m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/8S_jvsjtaYg"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240307-8s_jvsjtayg/"&gt;FAQ&lt;/a&gt; | 04 - XML intro and scraping (2h 3m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/I3auyTYORTs"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240220-i3auytyorts/"&gt;FAQ&lt;/a&gt; | 02 - Fundamentals of web scraping with urllib and BeautifulSoup (1h 37m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/DryMIxMf3VU"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240220-drymixmf3vu/"&gt;FAQ&lt;/a&gt; | 03 - Intermediate web scraping use of cookies (2h 19m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/cAriusuJsmw"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20240220-cariusujsmw/"&gt;FAQ&lt;/a&gt; | 01 - Intro to Web scraping and HTML (1h 33m)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://youtu.be/s9zNQoYNct8"&gt;Video&lt;/a&gt; | &lt;a href="/live-sessions/20230622-s9znqoynct8/"&gt;FAQ&lt;/a&gt; | Live Session - 2 (0h 29m)&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/llamafile/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llamafile/</guid><description>&lt;h2 id="local-llms-llamafile"&gt;Local LLMs: Llamafile&lt;a class="anchor" href="#local-llms-llamafile"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You would have heard of Large Language Models (LLMs) like GPT-4, Claude, and Llama. Some of these models are available for free, but most of them are not.&lt;/p&gt;
&lt;p&gt;An easy way to run LLMs locally is Mozilla&amp;rsquo;s &lt;a href="https://github.com/Mozilla-Ocho/llamafile"&gt;Llamafile&lt;/a&gt;. It&amp;rsquo;s a single executable file that works on Windows, Mac, and Linux. No installation or configuration needed - just download and run.&lt;/p&gt;
&lt;p&gt;Watch this Llamafile Tutorial (6 min):&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/d1Fnfvat6nM"&gt;&lt;img src="https://img.youtube.com/vi/d1Fnfvat6nM/0.jpg" alt="Llamafile: Local LLMs Made Easy" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-agents/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-agents/</guid><description>&lt;h2 id="llm-agents-building-ai-systems-that-can-think-and-act"&gt;LLM Agents: Building AI Systems That Can Think and Act&lt;a class="anchor" href="#llm-agents-building-ai-systems-that-can-think-and-act"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;LLM Agents are AI systems that can define and execute their own workflows to accomplish tasks. Unlike simple prompt-response patterns, agents make multiple LLM calls, use tools, and adapt their approach based on intermediate results. They represent a significant step toward more autonomous AI systems.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/DWUdGhRrv2c"&gt;&lt;img src="https://i.ytimg.com/vi_webp/DWUdGhRrv2c/sddefault.webp" alt="Building LLM Agents with LangChain (13 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="what-makes-an-agent"&gt;What Makes an Agent?&lt;a class="anchor" href="#what-makes-an-agent"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;An LLM agent consists of three core components:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-evals/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-evals/</guid><description>&lt;h2 id="llm-evaluations-with-promptfoo"&gt;LLM Evaluations with PromptFoo&lt;a class="anchor" href="#llm-evaluations-with-promptfoo"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Test-drive your prompts and models with automated, reliable evaluations.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/KhINc5XwhKs"&gt;&lt;img src="https://i.ytimg.com/vi_webp/KhINc5XwhKs/sddefault.webp" alt="🚀 Test Driven Prompt Engineering with PromptFoo (12 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;PromptFoo is a test-driven development framework for LLMs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Developer-first&lt;/strong&gt;: Fast CLI with live reload &amp;amp; caching (&lt;a href="https://promptfoo.dev"&gt;promptfoo.dev&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi-provider&lt;/strong&gt;: Works with OpenAI, Anthropic, HuggingFace, Ollama &amp;amp; more (&lt;a href="https://github.com/promptfoo/promptfoo"&gt;GitHub&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assertions&lt;/strong&gt;: Built‑in (&lt;code&gt;contains&lt;/code&gt;, &lt;code&gt;equals&lt;/code&gt;) &amp;amp; model‑graded (&lt;code&gt;llm-rubric&lt;/code&gt;) (&lt;a href="https://www.promptfoo.dev/docs/configuration/expected-outputs/"&gt;docs&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CI/CD&lt;/strong&gt;: Integrate evals into pipelines for regression safety (&lt;a href="https://www.promptfoo.dev/docs/integrations/ci-cd/"&gt;CI/CD guide&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To run PromptFoo:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-image-generation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-image-generation/</guid><description>&lt;h2 id="gemini-flash-experimental-image-generation-and-editing-apis"&gt;Gemini Flash Experimental Image Generation and Editing APIs&lt;a class="anchor" href="#gemini-flash-experimental-image-generation-and-editing-apis"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In March 2025, Google introduced native image generation and editing capabilities in the Gemini 2.0 Flash Experimental model. You can now generate and iteratively edit images via a single REST endpoint (&lt;a href="https://developers.googleblog.com/en/experiment-with-gemini-20-flash-native-image-generation/"&gt;Experiment with Gemini 2.0 Flash native image generation&lt;/a&gt;, &lt;a href="https://ai.google.dev/gemini-api/docs/image-generation"&gt;Generate images | Gemini API | Google AI for Developers&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/wgs4UYx6quY"&gt;&lt;img src="https://i.ytimg.com/vi_webp/wgs4UYx6quY/sddefault.webp" alt="How to use Latest Gemini 2.0 Native Image Generation with API? (9 min)" /&gt;&lt;/a&gt; (&lt;a href="https://www.youtube.com/watch?v=wgs4UYx6quY"&gt;How to use Latest Gemini 2.0 Native Image Generation with API?&lt;/a&gt;)&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-realtime/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-realtime/</guid><description/></item><item><title/><link>/2026-05/llm-sentiment-analysis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-sentiment-analysis/</guid><description>&lt;h2 id="llm-sentiment-analysis"&gt;LLM Sentiment Analysis&lt;a class="anchor" href="#llm-sentiment-analysis"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://platform.openai.com/"&gt;OpenAI&amp;rsquo;s API&lt;/a&gt; provides access to language models like GPT 4o, GPT 4o mini, etc.&lt;/p&gt;
&lt;p&gt;For more details, read OpenAI&amp;rsquo;s guide for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/text-generation"&gt;Text Generation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/vision"&gt;Vision&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/structured-outputs"&gt;Structured Outputs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Start with this quick tutorial:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/Xz4ORA0cOwQ"&gt;&lt;img src="https://i.ytimg.com/vi_webp/Xz4ORA0cOwQ/sddefault.webp" alt="OpenAI API Quickstart: Send your First API Request" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal example using &lt;code&gt;curl&lt;/code&gt; to generate text:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl https://api.openai.com/v1/chat/completions &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -H &lt;span style="color:#e6db74"&gt;&amp;#34;Content-Type: application/json&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -H &lt;span style="color:#e6db74"&gt;&amp;#34;Authorization: Bearer &lt;/span&gt;$OPENAI_API_KEY&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -d &lt;span style="color:#e6db74"&gt;&amp;#39;{
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;model&amp;#34;: &amp;#34;gpt-4o-mini&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;messages&amp;#34;: [{ &amp;#34;role&amp;#34;: &amp;#34;user&amp;#34;, &amp;#34;content&amp;#34;: &amp;#34;Write a haiku about programming.&amp;#34; }]
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; }&amp;#39;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Let&amp;rsquo;s break down the request:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-speech/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-speech/</guid><description>&lt;h2 id="openai-tts-1-for-text-to-speech-generation"&gt;OpenAI TTS-1 for Text-to-Speech Generation&lt;a class="anchor" href="#openai-tts-1-for-text-to-speech-generation"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;OpenAI&amp;rsquo;s Text-to-Speech API (TTS-1) converts text into natural-sounding speech using state-of-the-art neural models. Released in March 2025, it offers multiple voices and control over speaking style and speed.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/lXb0L16ISAc"&gt;&lt;img src="https://i.ytimg.com/vi_webp/lXb0L16ISAc/sddefault.webp" alt="Audio Models in the API (15 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="simple-speech-generation"&gt;Simple speech generation&lt;a class="anchor" href="#simple-speech-generation"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;To generate speech from text, send a POST request to the speech endpoint:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl https://api.openai.com/v1/audio/speech &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -H &lt;span style="color:#e6db74"&gt;&amp;#34;Authorization: Bearer &lt;/span&gt;$OPENAI_API_KEY&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -H &lt;span style="color:#e6db74"&gt;&amp;#34;Content-Type: application/json&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -d &lt;span style="color:#e6db74"&gt;&amp;#39;{
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;model&amp;#34;: &amp;#34;tts-1&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;input&amp;#34;: &amp;#34;Hello! This is a test of the OpenAI text to speech API.&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;voice&amp;#34;: &amp;#34;alloy&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; }&amp;#39;&lt;/span&gt; --output speech.mp3&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="generation-options"&gt;Generation options&lt;a class="anchor" href="#generation-options"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Control the output with these parameters:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-text-extraction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-text-extraction/</guid><description>&lt;h2 id="llm-text-extraction"&gt;LLM Text Extraction&lt;a class="anchor" href="#llm-text-extraction"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="/2026-05/json/"&gt;JSON&lt;/a&gt; is one of the most widely used formats in the world for applications to exchange data.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/72514uGffPE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/72514uGffPE/sddefault.webp" alt="LLM Extraction" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This video explains how to use LLMs to extract structure from unstructured data, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LLM for Data Extraction&lt;/strong&gt;: Use OpenAI&amp;rsquo;s API to extract structured information from unstructured data like addresses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;JSON Schema&lt;/strong&gt;: Define a JSON schema to ensure consistent and structured output from the LLM.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prompt Engineering&lt;/strong&gt;: Craft effective prompts to guide the LLM&amp;rsquo;s response and improve accuracy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Cleaning&lt;/strong&gt;: Use string functions and OpenAI&amp;rsquo;s API to clean and standardize data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Analysis&lt;/strong&gt;: Analyze extracted data using Pandas to gain insights.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM Limitations&lt;/strong&gt;: Understand the limitations of LLMs, including potential errors and inconsistencies in output.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Production Use Cases&lt;/strong&gt;: Explore real-world applications of LLMs for data extraction, such as customer service email analysis.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-video-screen-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-video-screen-scraping/</guid><description>&lt;h2 id="llm-video-screen-scraping"&gt;LLM Video Screen-Scraping&lt;a class="anchor" href="#llm-video-screen-scraping"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Video screen-scraping with LLMs is a powerful technique for extracting structured data from screen recordings. This approach works with any visible screen content and bypasses traditional web scraping limitations like authentication or anti-scraping measures.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/2G1LqS6qO5s"&gt;&lt;img src="https://i.ytimg.com/vi_webp/2G1LqS6qO5s/sddefault.webp" alt="Screen Scraping with Gemini" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Key benefits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No setup cost or authentication handling&lt;/li&gt;
&lt;li&gt;Works with any visible screen content&lt;/li&gt;
&lt;li&gt;Full control over data exposure&lt;/li&gt;
&lt;li&gt;Extremely cost-effective (&amp;lt; $0.001 per short video)&lt;/li&gt;
&lt;li&gt;Bypasses anti-scraping measures&lt;/li&gt;
&lt;li&gt;Handles varying formats and layouts&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="quick-start-example"&gt;Quick Start Example&lt;a class="anchor" href="#quick-start-example"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Here&amp;rsquo;s a basic workflow using Google&amp;rsquo;s AI Studio and Gemini:&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm-website-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm-website-scraping/</guid><description>&lt;h1 id="llm-website-scraping"&gt;LLM Website Scraping&lt;a class="anchor" href="#llm-website-scraping"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;LLM-powered website scraping represents a paradigm shift from traditional web scraping. Instead of writing explicit code to navigate DOM structures and CSS selectors, you describe what data you want and let the LLM figure out how to extract it.&lt;/p&gt;
&lt;p&gt;This approach works for both simple public pages and complex scenarios requiring authentication, JavaScript rendering, and dynamic interactions. LLMs can write scraping code, control browsers, adapt to layout changes, and even handle captchas through browser automation.&lt;/p&gt;</description></item><item><title/><link>/2026-05/llm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/llm/</guid><description>&lt;h2 id="llm-cli-llm"&gt;LLM CLI: llm&lt;a class="anchor" href="#llm-cli-llm"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://pypi.org/project/llm"&gt;&lt;code&gt;llm&lt;/code&gt;&lt;/a&gt; is a command-line utility for interacting with large language models—simplifying prompts, managing models and plugins, logging every conversation, and extracting structured data for pipelines.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/QUXQNi6jQ30?t=100"&gt;&lt;img src="https://i.ytimg.com/vi_webp/QUXQNi6jQ30/sddefault.webp" alt="Language models on the command-line w/ Simon Willison" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="basic-usage"&gt;Basic Usage&lt;a class="anchor" href="#basic-usage"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://github.com/simonw/llm#installation"&gt;Install llm&lt;/a&gt;. Then set up your &lt;a href="https://platform.openai.com/api-keys"&gt;&lt;code&gt;OPENAI_API_KEY&lt;/code&gt;&lt;/a&gt; environment variable. See &lt;a href="https://github.com/simonw/llm?tab=readme-ov-file#getting-started"&gt;Getting started&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TDS Students&lt;/strong&gt;: See &lt;a href="/2026-05/large-language-models/"&gt;Large Language Models&lt;/a&gt; for instructions on how to get and use &lt;code&gt;OPENAI_API_KEY&lt;/code&gt;.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Run a simple prompt&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm &lt;span style="color:#e6db74"&gt;&amp;#39;five great names for a pet pelican&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Continue a conversation&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -c &lt;span style="color:#e6db74"&gt;&amp;#39;now do walruses&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Start a memory-aware chat session&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm chat
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Specify a model&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -m gpt-4.1-nano &lt;span style="color:#e6db74"&gt;&amp;#39;Summarize tomorrow’s meeting agenda&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Extract JSON output&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm &lt;span style="color:#e6db74"&gt;&amp;#39;List the top 5 Python viz libraries with descriptions&amp;#39;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; --schema-multi &lt;span style="color:#e6db74"&gt;&amp;#39;name,description&amp;#39;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Or use llm without installation using &lt;a href="/2026-05/uv/"&gt;&lt;code&gt;uvx&lt;/code&gt;&lt;/a&gt;:&lt;/p&gt;</description></item><item><title/><link>/2026-05/making-open-data-useful/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/making-open-data-useful/</guid><description>&lt;h1 id="making-open-data-useful-lessons-from-diagram-chasing"&gt;Making Open Data Useful: Lessons from Diagram Chasing&lt;a class="anchor" href="#making-open-data-useful-lessons-from-diagram-chasing"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Many datasets are technically available but practically unusable. Your job is to make them useful for people who are not data scientists. Design for the lowest friction, and let people share one precise slice of truth with a link.&lt;/p&gt;
&lt;p&gt;See how Diagram Chasing does this across projects like &lt;strong&gt;CBFC Watch&lt;/strong&gt;, &lt;strong&gt;Time Use Explorer&lt;/strong&gt;, &lt;strong&gt;BLR Water Log&lt;/strong&gt;, &lt;strong&gt;Votes in a name?&lt;/strong&gt;, and &lt;strong&gt;Who is my neta?&lt;/strong&gt; (&lt;a href="https://cbfc.watch/" title="CBFC Watch: Home"&gt;CBFC Watch&lt;/a&gt;)&lt;/p&gt;</description></item><item><title/><link>/2026-05/marimo/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/marimo/</guid><description>&lt;h2 id="interactive-notebooks-marimo"&gt;Interactive Notebooks: Marimo&lt;a class="anchor" href="#interactive-notebooks-marimo"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://marimo.app/"&gt;Marimo&lt;/a&gt; is a new take on notebooks that solves some headaches of Jupyter. It runs cells reactively - when you change one cell, all dependent cells update automatically, just like a spreadsheet.&lt;/p&gt;
&lt;p&gt;Marimo&amp;rsquo;s cells can&amp;rsquo;t be run out of order. This makes Marimo more reproducible and easier to debug, but requires a mental shift from the Jupyter/Colab way of working.&lt;/p&gt;
&lt;p&gt;It also runs Python directly in the browser and is quite interactive. &lt;a href="https://marimo.io/gallery"&gt;Browse the gallery of examples&lt;/a&gt;. With a wide variety of interactive widgets, It&amp;rsquo;s growing popular as an alternative to Streamlit for building data science web apps.&lt;/p&gt;</description></item><item><title/><link>/2026-05/markdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/markdown/</guid><description>&lt;h2 id="documentation-markdown"&gt;Documentation: Markdown&lt;a class="anchor" href="#documentation-markdown"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Markdown is a lightweight markup language for creating formatted text using a plain-text editor. It&amp;rsquo;s the standard for documentation in software projects and data science notebooks.&lt;/p&gt;
&lt;p&gt;Watch this introduction to Markdown (19 min):&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/HUBNt18RFbo"&gt;&lt;img src="https://i.ytimg.com/vi_webp/HUBNt18RFbo/sddefault.webp" alt="Markdown Crash Course (19 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Common Markdown syntax:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;# Heading 1
## Heading 2

**bold** and *italic*

- Bullet point
- Another point
 - Nested point

1. Numbered list
2. Second item

[Link text](https://url.com)
![Image alt](image.jpg)

```python
# Code block
def hello():
 print(&amp;#34;Hello&amp;#34;)
```

&amp;gt; Blockquote&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;There is also a &lt;a href="https://github.github.com/gfm/"&gt;GitHub Flavored Markdown&lt;/a&gt; standard which is popular. This includes extensions like:&lt;/p&gt;</description></item><item><title/><link>/2026-05/marks-dashboard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/marks-dashboard/</guid><description>&lt;h1 id="marks-dashboard"&gt;Marks Dashboard&lt;a class="anchor" href="#marks-dashboard"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="t---score-calculation"&gt;T - Score Calculation&lt;a class="anchor" href="#t---score-calculation"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The final score for the course is calculated as follows:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Components (All out of 100)&lt;/th&gt;
					&lt;th&gt;Weight&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Grade Assessments &lt;strong&gt;(GA)&lt;/strong&gt; &lt;br&gt; Best 7/9&lt;/td&gt;
					&lt;td&gt;20%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Project 1 &lt;strong&gt;(P1)&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;20%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Project 2 &lt;strong&gt;(P2)&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;20%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Remote Online Exam &lt;strong&gt;(ROE)&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;20%&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;End Term &lt;strong&gt;(ET)&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;20%&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h4 id="end-term-eligibility-best-45-first-gas-average--40"&gt;End Term Eligibility: Best 4/5 first GAs Average &amp;gt;= 40&lt;a class="anchor" href="#end-term-eligibility-best-45-first-gas-average--40"&gt;#&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;T = 0.2 GA + 0.2 P1 + 0.2 P2 + 0.2 ROE + 0.2 ET&lt;/p&gt;</description></item><item><title/><link>/2026-05/marp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/marp/</guid><description>&lt;h2 id="marp-markdown-presentation-ecosystem"&gt;Marp: Markdown Presentation Ecosystem&lt;a class="anchor" href="#marp-markdown-presentation-ecosystem"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/EzQ-p41wNEE"&gt;&lt;img src="https://i.ytimg.com/vi/EzQ-p41wNEE/sddefault.jpg" alt="Never use PowerPoint again (20 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://marp.app/"&gt;Marp&lt;/a&gt; (Markdown Presentation) is a powerful tool for creating presentations using Markdown. It converts Markdown files into slideshows, making it ideal for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Technical Presentations&lt;/strong&gt;: Code snippets, diagrams, and technical content&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Documentation&lt;/strong&gt;: Creating slide decks from existing documentation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Academic Slides&lt;/strong&gt;: Research presentations with math equations and citations&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Conference Talks&lt;/strong&gt;: Professional presentations with custom themes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Teaching Materials&lt;/strong&gt;: Educational content with rich formatting&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GitHub Pages&lt;/strong&gt;: Hosting presentations on the web&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This tutorial covers using Marp in VS Code to create professional presentations.&lt;/p&gt;</description></item><item><title/><link>/2026-05/multimodal-embeddings/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/multimodal-embeddings/</guid><description>&lt;h2 id="multimodal-embeddings"&gt;Multimodal Embeddings&lt;a class="anchor" href="#multimodal-embeddings"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Multimodal embeddings map &lt;strong&gt;text&lt;/strong&gt; and &lt;strong&gt;images&lt;/strong&gt; into the &lt;strong&gt;same&lt;/strong&gt; vector space, enabling direct comparison between, say, a caption — “A cute cat” — and an image of that cat. This unified representation powers real-world applications like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cross-modal search&lt;/strong&gt; (e.g. “find images of a sunset” via text queries)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content recommendation&lt;/strong&gt; (suggesting visually similar products to text descriptions)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clustering &amp;amp; retrieval&lt;/strong&gt; (grouping documents and their associated graphics)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anomaly detection&lt;/strong&gt; (spotting unusual image–text pairings)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;By reducing different data types to a common numeric form, you unlock richer search, enhanced recommendations, and tighter integration of visual and textual data.&lt;/p&gt;</description></item><item><title/><link>/2026-05/narratives-with-comics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/narratives-with-comics/</guid><description>&lt;h2 id="narratives-with-comics"&gt;Narratives with Comics&lt;a class="anchor" href="#narratives-with-comics"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/HZDqCQBpHGI"&gt;&lt;img src="https://i.ytimg.com/vi/HZDqCQBpHGI/sddefault.jpg" alt="Comic narratives with Google Sheets &amp;amp; Comicgen" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.google.com/spreadsheets/d/1b0DOfJnnx6MFcN955YqRqYafLb8XrH-zqtLaK2h5kkc/edit#gid=1534638946"&gt;Sample sheet&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/narratives-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/narratives-with-excel/</guid><description>&lt;h2 id="narratives-with-excel"&gt;Narratives with Excel&lt;a class="anchor" href="#narratives-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/CRNJerr3pI4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/CRNJerr3pI4/sddefault.webp" alt="Narratives with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.google.com/spreadsheets/d/1Htmr5ar0ZX2nYW8Xerr9OaLJl3zS0gxY/view#gid=171350107"&gt;Sample Excel&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/network-analysis-in-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/network-analysis-in-python/</guid><description>&lt;h2 id="network-analysis-in-python"&gt;Network Analysis in Python&lt;a class="anchor" href="#network-analysis-in-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/uPL3VuRqOy4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/uPL3VuRqOy4/sddefault.webp" alt="Talk: Exploring the Movie Actor Network in Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to use network analysis to identify clusters and connections between nodes in a dataset, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Network Construction&lt;/strong&gt;: Build a network from the IMDb database, where nodes represent actors and edges represent shared movie appearances.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clustering&lt;/strong&gt;: Apply clustering techniques to detect communities within the network, using the scikit-network library.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Matrix Operations&lt;/strong&gt;: Utilize matrix operations to efficiently analyze actor relationships and interactions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community Detection&lt;/strong&gt;: Implement algorithms to identify and interpret clusters, examining how different actor clusters are connected.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Application of Findings&lt;/strong&gt;: Explore practical applications of network analysis, such as social network analysis and its potential uses in various domains.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/ngrok/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ngrok/</guid><description>&lt;h2 id="tunneling-ngrok"&gt;Tunneling: ngrok&lt;a class="anchor" href="#tunneling-ngrok"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://ngrok.com/"&gt;Ngrok&lt;/a&gt; is a tool that creates secure tunnels to your localhost, making your local development server accessible to the internet. It&amp;rsquo;s essential for testing webhooks, sharing work in progress, or debugging applications in production-like environments.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/dfMdLGZLXSg"&gt;&lt;img src="https://i.ytimg.com/vi_webp/dfMdLGZLXSg/sddefault.webp" alt="ngrok in 60 seconds" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Run the command &lt;code&gt;uvx ngrok http 8000&lt;/code&gt; to create a tunnel to your local server on port 8000. This generates a public URL that you can share with others.&lt;/p&gt;</description></item><item><title/><link>/2026-05/nominatim-api-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/nominatim-api-with-python/</guid><description>&lt;h2 id="nominatim-api-with-python"&gt;Nominatim API with Python&lt;a class="anchor" href="#nominatim-api-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/f0PZ-pphAXE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/f0PZ-pphAXE/sddefault.webp" alt="Nominatim Open Street Map with Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to get the latitude and longitude of any city from the Nominatim API.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Introduction to Nominatim&lt;/strong&gt;: Understand how Nominatim, from OpenStreetMap, works similarly to Google Maps for geocoding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Installation and Import&lt;/strong&gt;: Learn to install and import &lt;a href="https://geopy.readthedocs.io/"&gt;geopy&lt;/a&gt; and &lt;a href="https://nominatim.org/"&gt;nominatim&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Using the Locator&lt;/strong&gt;: Create a locator object using Nominatim and set up a user agent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Geocoding an Address&lt;/strong&gt;: Use &lt;code&gt;locator.geocode&lt;/code&gt; to input an address (e.g., Eiffel Tower) and fetch geocoded data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracting Data&lt;/strong&gt;: Access detailed information like latitude, longitude, bounding box, and accurate address from the JSON response.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Classifying Locations&lt;/strong&gt;: Identify the type of place (e.g., tourism, university) using the response data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Practical Example&lt;/strong&gt;: Geocode &amp;ldquo;IIT Madras&amp;rdquo; and retrieve its full address, type (university), and other relevant information.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links and references:&lt;/p&gt;</description></item><item><title/><link>/2026-05/npx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/npx/</guid><description>&lt;h2 id="javascript-tools-npx"&gt;JavaScript tools: npx&lt;a class="anchor" href="#javascript-tools-npx"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://docs.npmjs.com/cli/v8/commands/npx"&gt;npx&lt;/a&gt; is a command-line tool that comes with npm (Node Package Manager) and allows you to execute npm package binaries and run one-off commands without installing them globally. It&amp;rsquo;s essential for modern JavaScript development and data science workflows.&lt;/p&gt;
&lt;p&gt;For data scientists, npx is useful when:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Running JavaScript-based data visualization tools&lt;/li&gt;
&lt;li&gt;Converting notebooks and documents&lt;/li&gt;
&lt;li&gt;Testing and formatting code&lt;/li&gt;
&lt;li&gt;Running development servers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are common npx commands:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Run a package without installing&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx http-server . &lt;span style="color:#75715e"&gt;# Start a local web server&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx prettier --write . &lt;span style="color:#75715e"&gt;# Format code or docs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx eslint . &lt;span style="color:#75715e"&gt;# Lint JavaScript&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx ts-node script.ts &lt;span style="color:#75715e"&gt;# Run TypeScript directly&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx esbuild app.js &lt;span style="color:#75715e"&gt;# Bundle JavaScript&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx jsdoc . &lt;span style="color:#75715e"&gt;# Generate JavaScript docs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Run specific versions&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx prettier@3.6 --write . &lt;span style="color:#75715e"&gt;# Use prettier 3.6&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Execute remote scripts (use with caution!)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;npx github:user/repo &lt;span style="color:#75715e"&gt;# Run from GitHub&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Watch this introduction to npx (6 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/ollama/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/ollama/</guid><description>&lt;h2 id="local-llm-runner-ollama"&gt;Local LLM Runner: Ollama&lt;a class="anchor" href="#local-llm-runner-ollama"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/ollama/ollama"&gt;&lt;code&gt;ollama&lt;/code&gt;&lt;/a&gt; is a command-line tool for running open-source large language models entirely on your own machine—no API keys, no vendor lock-in, full control over models and performance.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/Lb5D892-2HY"&gt;&lt;img src="https://i.ytimg.com/vi_webp/Lb5D892-2HY/sddefault.webp" alt="Run AI Models Locally: Ollama Tutorial (Step-by-Step Guide &amp;#43; WebUI)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="basic-usage"&gt;Basic Usage&lt;a class="anchor" href="#basic-usage"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://ollama.com/"&gt;Download Ollama for macOS, Linux, or Windows&lt;/a&gt; and add the binary to your &lt;code&gt;PATH&lt;/code&gt;. See the full &lt;a href="https://ollama.com/docs"&gt;Docs ↗&lt;/a&gt; for installation details and troubleshooting.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# List installed and available models&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ollama list
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Download/pin a specific model version&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ollama pull gemma3:1b-it-qat
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Run a one-off prompt&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ollama run gemma3:1b-it-qat &lt;span style="color:#e6db74"&gt;&amp;#39;Write a haiku about data visualization&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Launch a persistent HTTP API on port 11434&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ollama serve
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Interact programmatically over HTTP&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl -X POST http://localhost:11434/api/chat &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -H &lt;span style="color:#e6db74"&gt;&amp;#39;Content-Type: application/json&amp;#39;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; -d &lt;span style="color:#e6db74"&gt;&amp;#39;{&amp;#34;model&amp;#34;:&amp;#34;gemma3:1b-it-qat&amp;#34;,&amp;#34;prompt&amp;#34;:&amp;#34;Hello, world!&amp;#34;}&amp;#39;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="key-features"&gt;Key Features&lt;a class="anchor" href="#key-features"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model management&lt;/strong&gt;: &lt;code&gt;list&lt;/code&gt;/&lt;code&gt;pull&lt;/code&gt; — Install and switch among Llama 3.3, DeepSeek-R1, Gemma 3, Mistral, Phi-4, and more.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Local inference&lt;/strong&gt;: &lt;code&gt;run&lt;/code&gt; — Execute prompts entirely on-device for privacy and zero latency beyond hardware limits.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Persistent server&lt;/strong&gt;: &lt;code&gt;serve&lt;/code&gt; — Expose a local REST API for multi-session chats and integration into scripts or apps.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Version pinning&lt;/strong&gt;: &lt;code&gt;pull model:tag&lt;/code&gt; — Pin exact model versions for reproducible demos and experiments.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resource control&lt;/strong&gt;: &lt;code&gt;--threads&lt;/code&gt; / &lt;code&gt;--context&lt;/code&gt; — Tune CPU/GPU usage and maximum context window for performance and memory management.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="real-world-use-cases"&gt;Real-World Use Cases&lt;a class="anchor" href="#real-world-use-cases"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Quick prototyping&lt;/strong&gt;. Brainstorm slide decks or blog outlines offline, without worrying about API quotas: &lt;code&gt;ollama run gemma-3 'Outline a slide deck on Agile best practices'&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data privacy&lt;/strong&gt;. Summarize sensitive documents on-device, retaining full control of your data: &lt;code&gt;cat financial_report.pdf | ollama run phi-4 'Summarize the key findings'&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CI/CD integration&lt;/strong&gt;. Validate PR descriptions or test YAML configurations in your pipeline without incurring API costs: &lt;code&gt;git diff origin/main | ollama run llama2 'Check for style and clarity issues'&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Local app embedding&lt;/strong&gt;. Power a desktop or web app via the local REST API for instant LLM features: &lt;code&gt;curl -X POST http://localhost:11434/api/chat -d '{&amp;quot;model&amp;quot;:&amp;quot;mistral&amp;quot;,&amp;quot;prompt&amp;quot;:&amp;quot;Translate to German&amp;quot;}'&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Read the full &lt;a href="https://github.com/ollama/ollama/tree/main/docs"&gt;Ollama docs ↗&lt;/a&gt; for advanced topics like custom model hosting, GPU tuning, and integrating with your development workflows.&lt;/p&gt;</description></item><item><title/><link>/2026-05/outlier-detection-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/outlier-detection-with-excel/</guid><description>&lt;h2 id="outlier-detection-with-excel"&gt;Outlier Detection with Excel&lt;a class="anchor" href="#outlier-detection-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/sUTJb0F9eBw"&gt;&lt;img src="https://i.ytimg.com/vi_webp/sUTJb0F9eBw/sddefault.webp" alt="Outlier detection with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to identify and handle outliers in data using Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Understanding Outliers&lt;/strong&gt;: Definition of outliers and their impact on statistical analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Calculating Quartiles&lt;/strong&gt;: Using Excel formulas to calculate Q1 (first quartile) and Q3 (third quartile).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interquartile Range (IQR)&lt;/strong&gt;: Finding the IQR by subtracting Q1 from Q3.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Determining Bounds&lt;/strong&gt;: Calculating lower and upper bounds using 1.5 times the IQR.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Identifying Outliers&lt;/strong&gt;: Using Excel functions to determine if data points fall outside the calculated bounds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visualizing Data&lt;/strong&gt;: Creating box plots to visualize outliers and data distribution.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Outliers&lt;/strong&gt;: Deciding whether to exclude or keep outliers based on their impact on analysis.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/parsing-json/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/parsing-json/</guid><description>&lt;h2 id="parsing-json"&gt;Parsing JSON&lt;a class="anchor" href="#parsing-json"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;JSON is everywhere—APIs, logs, configuration files—and its nested or large structure can challenge memory and processing. In this tutorial, we&amp;rsquo;ll explore tools to flatten, stream, and query JSON data efficiently.&lt;/p&gt;
&lt;p&gt;For example, we&amp;rsquo;ll often need to process a multi-gigabyte log file from a web service where each record is a JSON object.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/1lxrb_ezP-g"&gt;&lt;img src="https://i.ytimg.com/vi/1lxrb_ezP-g/sddefault.jpg" alt="JSON Parsing in Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This requires us to handle complex nested structures, large files that don&amp;rsquo;t fit in memory, or extract specific fields. Here are the key tools and techniques for efficient JSON parsing:&lt;/p&gt;</description></item><item><title/><link>/2026-05/profiling-data-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/profiling-data-with-python/</guid><description>&lt;h2 id="profile-data-with-python"&gt;Profile Data with Python&lt;a class="anchor" href="#profile-data-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/kFVxdBhLa_A"&gt;&lt;img src="https://i.ytimg.com/vi_webp/kFVxdBhLa_A/sddefault.webp" alt="Discover the data profile with Python" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This session covers the use of the &lt;code&gt;pandas_profiling&lt;/code&gt; library for generating comprehensive data reports in Python:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Library Installation and Import&lt;/strong&gt;: Learn how to install and import the pandas_profiling library.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Profile Report Generation&lt;/strong&gt;: Generate an HTML report with a single line of code using ProfileReport.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Descriptive Statistics&lt;/strong&gt;: View detailed descriptive statistics such as variance, standard deviation, and kurtosis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outlier Detection&lt;/strong&gt;: Identify and analyze outliers within the dataset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Correlation Analysis&lt;/strong&gt;: Understand how variables are correlated with each other using visual representations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Missing Values&lt;/strong&gt;: Get insights on missing data and decide on imputation or removal strategies.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Initial Data Insights&lt;/strong&gt;: Use the report to gather early warnings and insights before starting the data cleaning and modeling process.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-1/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-1/</guid><description>&lt;h1 id="project-1---llm-based-automation-agent"&gt;Project 1 - LLM-based Automation Agent&lt;a class="anchor" href="#project-1---llm-based-automation-agent"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This project is due on 16 Feb 2025 EoD IST. Results will be announced by 26 Feb 2025.&lt;/p&gt;
&lt;p&gt;For questions, &lt;a href="https://discourse.onlinedegree.iitm.ac.in/t/project-1-llm-based-automation-agent-discussion-thread-tds-jan-2025/164277"&gt;use this Discourse thread&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="background"&gt;Background&lt;a class="anchor" href="#background"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You have joined the operations team at &lt;strong&gt;DataWorks Solutions&lt;/strong&gt;, a company that processes large volumes of log files, reports, and code artifacts to generate actionable insights for internal stakeholders. In order to improve operational efficiency and consistency, the company has mandated that routine tasks be automated and integrated into their Continuous Integration (CI) pipeline.&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-2/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-2/</guid><description>&lt;h1 id="project-2---tds-solver"&gt;Project 2 - TDS Solver&lt;a class="anchor" href="#project-2---tds-solver"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This project is due on 31 Mar 2025 EoD IST. Results will be announced by 15 Apr 2025.&lt;/p&gt;
&lt;p&gt;For questions, &lt;a href="https://discourse.onlinedegree.iitm.ac.in/t/project-2-tds-solver-discussion-thread/169029"&gt;use this Discourse thread&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="background"&gt;Background&lt;a class="anchor" href="#background"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You are a clever student who has joined IIT Madras&amp;rsquo; Online Degree in Data Science. You have just enrolled in the &lt;a href="https://tds.s-anand.net/"&gt;Tools in Data Science&lt;/a&gt; course.&lt;/p&gt;
&lt;p&gt;To make your life easier, you have decided to build an LLM-based application that can automatically answer any of the graded assignment questions.&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-data-analyst-agent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-data-analyst-agent/</guid><description>&lt;h1 id="project-data-analyst-agent"&gt;Project: Data Analyst Agent&lt;a class="anchor" href="#project-data-analyst-agent"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Deploy a data analyst agent. This is an API that uses LLMs to source, prepare, analyze, and visualize any data.&lt;/p&gt;
&lt;p&gt;Your application exposes an API endpoint. You may host it anywhere. Let&amp;rsquo;s assume it&amp;rsquo;s at &lt;code&gt;https://app.example.com/api/&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The endpoint must accept a POST request, e.g. &lt;code&gt;POST https://app.example.com/api/&lt;/code&gt; with a data analysis task description and optional attachments in the body. For example:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl &lt;span style="color:#e6db74"&gt;&amp;#34;https://app.example.com/api/&amp;#34;&lt;/span&gt; -F &lt;span style="color:#e6db74"&gt;&amp;#34;questions.txt=@question.txt&amp;#34;&lt;/span&gt; -F &lt;span style="color:#e6db74"&gt;&amp;#34;image.png=@image.png&amp;#34;&lt;/span&gt; -F &lt;span style="color:#e6db74"&gt;&amp;#34;data.csv=@data.csv&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;questions.txt&lt;/code&gt; will ALWAYS be sent and contain the questions. There may be zero or more additional files passed.&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-llm-analysis-quiz/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-llm-analysis-quiz/</guid><description>&lt;h1 id="project-llm-analysis-quiz"&gt;Project: LLM Analysis Quiz&lt;a class="anchor" href="#project-llm-analysis-quiz"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;![WARNING] The API Project Endpoint will be at &lt;a href="https://tds-llm-analysis.s-anand.net/project2"&gt;https://tds-llm-analysis.s-anand.net/project2&lt;/a&gt;
This will be active from 3:00 pm - 4:00 pm IST on Sat 29 Nov 2025.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;In this project, students will build an application that can solve a quiz that involves data sourcing, preparation, analysis, and visualization using LLMs. You will also answer a viva about your design choices.&lt;/p&gt;
&lt;p&gt;Fill out this &lt;a href="https://forms.gle/V3vW2QeHGPF9BTrB7"&gt;Google Form&lt;/a&gt;. It asks for:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Your email address&lt;/li&gt;
&lt;li&gt;A secret string (used to verify your requests).&lt;/li&gt;
&lt;li&gt;A system prompt that resists revealing the given code word (which will be appended to the system prompt). Max 100 chars.&lt;/li&gt;
&lt;li&gt;A user prompt that will override any such system prompt to reveal the code word. Max 100 chars.&lt;/li&gt;
&lt;li&gt;Your API endpoint URL (where you will accept POST requests with quiz tasks). Prefer HTTPS.&lt;/li&gt;
&lt;li&gt;Your GitHub repo URL (where your code is hosted). Make sure it&amp;rsquo;s public and has an &lt;a href="https://docs.github.com/en/communities/setting-up-your-project-for-healthy-contributions/adding-a-license-to-a-repository"&gt;MIT LICENSE&lt;/a&gt; when we evaluate. You may keep it private during development.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="prompt-testing"&gt;Prompt Testing&lt;a class="anchor" href="#prompt-testing"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s how we will test your system and user prompts:&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-llm-code-deployment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-llm-code-deployment/</guid><description>&lt;h1 id="project-llm-code-deployment"&gt;Project: LLM Code Deployment&lt;a class="anchor" href="#project-llm-code-deployment"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;!-- https://chatgpt.com/c/68cb90cb-7a64-8331-bfbf-aa222df475de --&gt;
&lt;p&gt;In this project, students will build an application that can build, deploy, update an application!&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Build&lt;/strong&gt;. The student:
&lt;ul&gt;
&lt;li&gt;receives &amp;amp; verifies a &lt;strong&gt;request&lt;/strong&gt; containing an app brief&lt;/li&gt;
&lt;li&gt;uses an &lt;strong&gt;LLM-assisted generator&lt;/strong&gt; to build the app,&lt;/li&gt;
&lt;li&gt;deployes to &lt;strong&gt;GitHub Pages&lt;/strong&gt;,&lt;/li&gt;
&lt;li&gt;then &lt;strong&gt;pings an evaluation API&lt;/strong&gt; with repo details&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evaluate&lt;/strong&gt;. The instructors:
&lt;ul&gt;
&lt;li&gt;run automated &lt;strong&gt;static, dynamic (Playwright), and and LLM&lt;/strong&gt; checks&lt;/li&gt;
&lt;li&gt;store and publish the results &lt;strong&gt;after the deadline&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;send a &lt;strong&gt;second request&lt;/strong&gt; tailored to the student’s codebase&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Revise&lt;/strong&gt;. The student
&lt;ul&gt;
&lt;li&gt;verifies secret&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;updates&lt;/strong&gt; the app based on the request&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;re‑deploys&lt;/strong&gt; Pages&lt;/li&gt;
&lt;li&gt;then &lt;strong&gt;pings a second evaluation API&lt;/strong&gt; with repo metadata.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="request"&gt;Request&lt;a class="anchor" href="#request"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The request is a JSON file like this:&lt;/p&gt;</description></item><item><title/><link>/2026-05/project-tds-virtual-ta/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/project-tds-virtual-ta/</guid><description>&lt;h1 id="project-tds-virtual-ta"&gt;Project: TDS Virtual TA&lt;a class="anchor" href="#project-tds-virtual-ta"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Create a virtual Teaching Assistant Discourse responder.&lt;/p&gt;
&lt;h2 id="background"&gt;Background&lt;a class="anchor" href="#background"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You are a clever student who has joined IIT Madras&amp;rsquo; Online Degree in Data Science. You have just enrolled in the &lt;a href="https://tds.s-anand.net/#/2025-01/"&gt;Tools in Data Science&lt;/a&gt; course.&lt;/p&gt;
&lt;p&gt;Out of kindness for your teaching assistants, you have decided to build an API that can automatically answer student questions on their behalf based on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://tds.s-anand.net/#/2025-01/"&gt;Course content&lt;/a&gt; with content for TDS Jan 2025 as on 15 Apr 2025.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://discourse.onlinedegree.iitm.ac.in/c/courses/tds-kb/34"&gt;TDS Discourse posts&lt;/a&gt; with content from 1 Jan 2025 - 14 Apr 2025.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="scrape-the-data"&gt;Scrape the data&lt;a class="anchor" href="#scrape-the-data"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;To make sure you can answer these questions, you will need to extract the data from the above source.&lt;/p&gt;</description></item><item><title/><link>/2026-05/projects/01-project-1/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/projects/01-project-1/</guid><description>&lt;h1 id="project-1-after-week-3"&gt;Project 1 (After Week 3)&lt;a class="anchor" href="#project-1-after-week-3"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;P1 is available at &lt;a href="https://exam.sanand.workers.dev/tds-2026-05-p1"&gt;https://exam.sanand.workers.dev/tds-2026-05-p1&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/projects/02-project-2/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/projects/02-project-2/</guid><description>&lt;h1 id="project-2-after-week-6"&gt;Project 2 (After Week 6)&lt;a class="anchor" href="#project-2-after-week-6"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;P2 is available at &lt;a href="https://exam.sanand.workers.dev/tds-2026-05-p2"&gt;https://exam.sanand.workers.dev/tds-2026-05-p2&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;</description></item><item><title/><link>/2026-05/prompt-engineering/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/prompt-engineering/</guid><description>&lt;h2 id="prompt-engineering"&gt;Prompt Engineering&lt;a class="anchor" href="#prompt-engineering"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Prompt engineering is the process of crafting effective prompts for large language models (LLMs).&lt;/p&gt;
&lt;p&gt;One of the best ways to approach prompt engineering is to think of LLMs as a smart colleague (with amnesia) who needs explicit instructions.&lt;/p&gt;
&lt;p&gt;The most authoritative guides are from the LLM providers themselves:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/"&gt;Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/introduction-prompt-design"&gt;Google&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/prompt-engineering"&gt;OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are some best practices:&lt;/p&gt;
&lt;h3 id="use-prompt-optimizers"&gt;Use prompt optimizers&lt;a class="anchor" href="#use-prompt-optimizers"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;They rewrite your prompt to improve it. Explore:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/prompt-improver"&gt;Anthropic Prompt Optimizer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.openai.com/docs/guides/prompt-generation"&gt;OpenAI Prompt Generation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/ai-powered-prompt-writing"&gt;Google AI-powered prompt writing tools&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="be-clear-direct-and-detailed"&gt;Be clear, direct, and detailed&lt;a class="anchor" href="#be-clear-direct-and-detailed"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Be explicit and thorough. Include all necessary context, goals, and details so the model understands the full picture.&lt;/p&gt;</description></item><item><title/><link>/2026-05/pydantic-ai/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/pydantic-ai/</guid><description>&lt;h2 id="pydantic-ai"&gt;Pydantic AI&lt;a class="anchor" href="#pydantic-ai"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://docs.pydantic.dev/"&gt;Pydantic&lt;/a&gt; is Python&amp;rsquo;s most widely used data validation library. &lt;a href="https://ai.pydantic.dev/"&gt;Pydantic AI&lt;/a&gt; extends this to build production-grade AI agents with structured outputs and type safety.&lt;/p&gt;
&lt;h3 id="whats-pydantic"&gt;What&amp;rsquo;s Pydantic?&lt;a class="anchor" href="#whats-pydantic"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;Python doesn&amp;rsquo;t enforce type hints at runtime. Pydantic validates data automatically using type annotations:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# /// script&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# requires-python = &amp;#34;&amp;gt;=3.11&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# dependencies = [&amp;#34;pydantic&amp;#34;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# ///&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; pydantic &lt;span style="color:#f92672"&gt;import&lt;/span&gt; BaseModel, Field
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;class&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;User&lt;/span&gt;(BaseModel):
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; name: str
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; age: int &lt;span style="color:#f92672"&gt;=&lt;/span&gt; Field(&lt;span style="color:#f92672"&gt;...&lt;/span&gt;, gt&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;, le&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;120&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; email: str
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Automatic validation and type conversion&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;user &lt;span style="color:#f92672"&gt;=&lt;/span&gt; User(name&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Alice&amp;#34;&lt;/span&gt;, age&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;25&amp;#34;&lt;/span&gt;, email&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;alice@example.com&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(user&lt;span style="color:#f92672"&gt;.&lt;/span&gt;age, type(user&lt;span style="color:#f92672"&gt;.&lt;/span&gt;age)) &lt;span style="color:#75715e"&gt;# 25 &amp;lt;class &amp;#39;int&amp;#39;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Invalid data raises ValidationError&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;try&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; User(name&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Bob&amp;#34;&lt;/span&gt;, age&lt;span style="color:#f92672"&gt;=-&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;5&lt;/span&gt;, email&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;invalid&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;except&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;Exception&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; e:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Validation failed: &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;e&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Pydantic is used by FastAPI, LangChain, the OpenAI SDK, and 8,000+ packages on PyPI. It&amp;rsquo;s fast (core written in Rust), integrates with IDEs, and generates JSON schemas automatically.&lt;/p&gt;</description></item><item><title/><link>/2026-05/rag-cli/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/rag-cli/</guid><description>&lt;h2 id="retrieval-augmented-generation-rag-with-the-cli"&gt;Retrieval Augmented Generation (RAG) with the CLI&lt;a class="anchor" href="#retrieval-augmented-generation-rag-with-the-cli"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Retrieval Augmented Generation (RAG) combines retrieval (searching a knowledge base) with generation (using an LLM) to produce answers grounded in your own documents. Instead of relying solely on a general-purpose LLM, RAG lets you feed it the most relevant chunks from your corpus at query time, improving accuracy, reducing hallucinations, and allowing you to answer domain‑specific questions without fine‑tuning.&lt;/p&gt;
&lt;p&gt;In particular, you can answer questions that are hard to answer with a keyword search. For example:&lt;/p&gt;</description></item><item><title/><link>/2026-05/rawgraphs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/rawgraphs/</guid><description>&lt;h2 id="rawgraphs"&gt;RAWgraphs&lt;a class="anchor" href="#rawgraphs"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/2TtYlty-M5g"&gt;&lt;img src="https://i.ytimg.com/vi_webp/2TtYlty-M5g/sddefault.webp" alt="RAWGraphs 1.0 - Introduction (1 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.rawgraphs.io/"&gt;RAWgraphs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You&amp;rsquo;ll learn about RAWGraphs and how to create data visualizations, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Understanding RAWGraphs&lt;/strong&gt;: Learn that RAWGraphs is an &lt;strong&gt;open-source data visualization tool&lt;/strong&gt; designed to simplify the visualization of complex data for everyone.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Import&lt;/strong&gt;: Discover how to &lt;strong&gt;import your dataset&lt;/strong&gt; into RAWGraphs by copying and pasting it from any spreadsheet application.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Chart Selection&lt;/strong&gt;: Understand how to &lt;strong&gt;choose from a variety of charts&lt;/strong&gt; that are built on top of d3js, such as a scatter plot, suitable for your data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Chart Creation&lt;/strong&gt;: Create different kinds of charts, notably non-traditional charts that go beyond bars, lines, pie, etc.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mapping Data Dimensions&lt;/strong&gt;: Learn the process of &lt;strong&gt;mapping your dataset&amp;rsquo;s dimensions to the visual variables&lt;/strong&gt; of your chosen chart. This includes examples like mapping rating to the horizontal axis, production budget to the vertical axis, box office to bubble areas, genre to color, and movie names to labels.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Customization and Export&lt;/strong&gt;: Find out how to &lt;strong&gt;customize your visualization&lt;/strong&gt; and &lt;strong&gt;export it as a vector or raster image&lt;/strong&gt;, which can then be fine-tuned using any graphic editor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Further Resources&lt;/strong&gt;: Be informed about where to &lt;strong&gt;find more information&lt;/strong&gt; on RAWGraphs&amp;rsquo; website and how to &lt;strong&gt;contribute to the project&lt;/strong&gt; through GitHub.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Watch these videos to learn how to create non-traditional data visualizations with &lt;a href="https://rawgraphs.io/"&gt;RawGraphs&lt;/a&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/reference/01-tools-glossary/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/reference/01-tools-glossary/</guid><description>&lt;h1 id="tools-glossary"&gt;Tools Glossary&lt;a class="anchor" href="#tools-glossary"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Every tool and term used across the course, with one line on what it&amp;rsquo;s for and where it&amp;rsquo;s taught.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Use &lt;code&gt;Ctrl+F&lt;/code&gt;. Terms are grouped by what you&amp;rsquo;re trying to do, not alphabetically — you usually know the &lt;em&gt;job&lt;/em&gt;, not the name.&lt;/p&gt;
&lt;h2 id="getting-data-off-the-web"&gt;Getting data off the web&lt;a class="anchor" href="#getting-data-off-the-web"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Hidden / internal API&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The JSON endpoint a page&amp;rsquo;s own JavaScript calls. Almost always better than parsing HTML → &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;httpx&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Modern Python HTTP client; sync + async, HTTP/2, connection reuse&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;selectolax&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Very fast HTML parser; CSS selectors (a lighter alternative to BeautifulSoup)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Playwright&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Browser automation that renders JavaScript; auto-waits for elements&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Selenium&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The older browser-automation standard; manual waits&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Patchright&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Drop-in Playwright fork patched to hide automation signals&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Camoufox&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Anti-detect Firefox build with fingerprint spoofing&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;curl_cffi&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;HTTP client that impersonates a browser&amp;rsquo;s TLS/JA3 and HTTP/2 fingerprint&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;CDX API&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Internet Archive&amp;rsquo;s index API — lists every snapshot of a URL&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;WARC&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Web ARChive format; raw request/response records used by Common Crawl&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;JSON-LD&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Structured data (schema.org) embedded in a page for search engines&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Sitemap index&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A sitemap listing other sitemaps; needs one extra level of parsing&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;robots.txt&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A site&amp;rsquo;s machine-readable crawling preferences. Not a law; strong evidence → &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical&lt;/a&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Crawl-delay&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A &lt;code&gt;robots.txt&lt;/code&gt; directive asking for a minimum gap between requests&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="being-blocked-and-not-being-blocked"&gt;Being blocked (and not being blocked)&lt;a class="anchor" href="#being-blocked-and-not-being-blocked"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;JA3 / JA4&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Fingerprints of a TLS handshake; identifies your client before any header is read&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;HTTP/2 fingerprint&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Identification from SETTINGS frames and header ordering&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Turnstile&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Cloudflare&amp;rsquo;s CAPTCHA replacement; usually invisible&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Bot score&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Cloudflare&amp;rsquo;s 1–99 automation likelihood; 1 = certainly bot&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;cf_clearance&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Cookie proving a challenge was passed; bound to IP + User-Agent&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;WAF&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Web Application Firewall — rule-based request filtering&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Verified bot&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A known-good crawler (Googlebot, uptime monitors) that rules should exempt&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Exponential backoff&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Doubling the wait after each failure&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Jitter&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Randomness added to backoff so clients don&amp;rsquo;t retry in lockstep&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;Retry-After&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Header telling you exactly how long to wait after a 429&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Idempotent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Safe to run twice — the second run changes nothing&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="storing-and-querying"&gt;Storing and querying&lt;a class="anchor" href="#storing-and-querying"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;SQLite&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;In-process OLTP database; ideal for scraper state&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;DuckDB&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;In-process OLAP database; SQL over Parquet/CSV/JSON files&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Parquet&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Columnar, compressed file format; 5–10× smaller than CSV&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;OLTP / OLAP&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Many small transactions vs few large analytical scans&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Content hash&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Hash of meaningful fields, used to detect changes&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Stable ID&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;An identifier for a record that survives re-runs&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;polars&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Fast DataFrame library; a Pandas alternative&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Dead-letter queue&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Where messages go after repeated processing failures&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="documents-images-audio-video"&gt;Documents, images, audio, video&lt;a class="anchor" href="#documents-images-audio-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;trafilatura&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Extracts the main article content from a page (drops boilerplate)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;markdownify&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Converts HTML to Markdown verbatim&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;MarkItDown&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Microsoft converter: PDF/DOCX/PPTX/XLSX → Markdown&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Jina Reader&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Hosted &lt;code&gt;r.jina.ai/&amp;lt;url&amp;gt;&lt;/code&gt; → clean Markdown&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;PyMuPDF&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Fast PDF text and table extraction&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Docling&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Layout-aware document → structured Markdown&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Surya&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Open multilingual OCR and layout analysis&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;OCR&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Optical Character Recognition — text from images&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;EXIF&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Image metadata; can include GPS coordinates&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;VLM&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Vision-Language Model — reads images and answers in text&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Qwen3-VL&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Leading open-weight vision-language family (2026)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Whisper / faster-whisper&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Speech-to-text; &lt;code&gt;faster-whisper&lt;/code&gt; is the quicker runtime&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Parakeet&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;NVIDIA STT; fastest self-hosted English throughput&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;WhisperX&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Whisper plus word-level alignment and speaker diarization&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Diarization&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Working out &lt;em&gt;who&lt;/em&gt; spoke &lt;em&gt;when&lt;/em&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;WER&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Word Error Rate — the standard STT accuracy metric&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;ffmpeg / ffprobe&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Convert/extract media; inspect media metadata&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="recon-and-osint"&gt;Recon and OSINT&lt;a class="anchor" href="#recon-and-osint"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;OSINT&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Open-Source Intelligence — findings assembled from public sources&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Dorking&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Using advanced search operators to find specific content&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;GHDB&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Google Hacking Database — catalogue of exposure patterns&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Certificate Transparency&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Public append-only logs of issued TLS certificates&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;crt.sh&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Search interface over CT logs&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;RDAP&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The modern, structured replacement for WHOIS&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Subdomain enumeration&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Discovering hostnames for a domain (amass, subfinder)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Shodan / Censys&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Search engines for internet-facing services&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Corroboration&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Confirming a claim across independent sources before asserting it&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Provenance&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The record of where a claim came from and when&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Aggregation harm&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Individually-public facts becoming harmful once combined&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Gitleaks / TruffleHog&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Secret scanners for git history&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="security-cloud-cicd"&gt;Security, cloud, CI/CD&lt;a class="anchor" href="#security-cloud-cicd"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Prompt injection&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Untrusted text being treated as instructions by a model&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Indirect injection&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Injection delivered through content the model &lt;em&gt;reads&lt;/em&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;OWASP LLM Top 10&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The reference list of LLM application risks&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Least privilege&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Granting only the permissions strictly required&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Multi-stage build&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A Dockerfile pattern keeping build tools out of the final image&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Layer cache&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Docker&amp;rsquo;s reuse of unchanged build steps&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Cold start&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;First-request latency after scaling to zero&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Scale to zero&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Running no instances (and paying nothing) when idle&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;IaC&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Infrastructure as Code — declared in version-controlled files&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Terraform state&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Terraform&amp;rsquo;s record of what actually exists; never commit it&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Drift&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Real infrastructure diverging from the declared config&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;WIF&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Workload Identity Federation — cloud auth without long-lived keys&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;At-least-once delivery&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A message may arrive more than once; consumers must be idempotent&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Pub/Sub&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Publish–subscribe messaging; decouples producers from consumers&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="course-wide"&gt;Course-wide&lt;a class="anchor" href="#course-wide"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;What it is&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;uv&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Fast Python package/project manager used throughout&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;PEP 723&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Inline script metadata — dependencies declared in the file itself&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;RAG&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Retrieval-Augmented Generation&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;MCP&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Model Context Protocol&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;OIDC&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;OpenID Connect&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Structured output&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Forcing a model to return schema-valid JSON&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;</description></item><item><title/><link>/2026-05/reference/02-command-cheatsheet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/reference/02-command-cheatsheet/</guid><description>&lt;h1 id="command-cheatsheet"&gt;Command Cheatsheet&lt;a class="anchor" href="#command-cheatsheet"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;The commands you&amp;rsquo;ll actually retype. Copy, adapt, move on.&lt;/p&gt;
&lt;/blockquote&gt;&lt;h2 id="uv--python-without-the-ceremony"&gt;uv — Python, without the ceremony&lt;a class="anchor" href="#uv--python-without-the-ceremony"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv run script.py &lt;span style="color:#75715e"&gt;# run a PEP 723 script; deps install automatically&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv run --with httpx script.py &lt;span style="color:#75715e"&gt;# add a one-off dependency&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uvx ruff check . &lt;span style="color:#75715e"&gt;# run a tool without installing it&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uvx duckdb -c &lt;span style="color:#e6db74"&gt;&amp;#34;SELECT 42&amp;#34;&lt;/span&gt; &lt;span style="color:#75715e"&gt;# same, for DuckDB&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv init myproject &lt;span style="color:#f92672"&gt;&amp;amp;&amp;amp;&lt;/span&gt; cd myproject &lt;span style="color:#75715e"&gt;# new project&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv add httpx polars duckdb &lt;span style="color:#75715e"&gt;# add dependencies&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv sync --frozen &lt;span style="color:#75715e"&gt;# install exactly the lockfile (use in CI)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv python install 3.13 &lt;span style="color:#75715e"&gt;# install a Python version&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The PEP 723 header that makes a single file runnable anywhere:&lt;/p&gt;</description></item><item><title/><link>/2026-05/regression-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/regression-with-excel/</guid><description>&lt;h2 id="regression-with-excel"&gt;Regression with Excel&lt;a class="anchor" href="#regression-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/AERQBMIHwXA"&gt;&lt;img src="https://i.ytimg.com/vi_webp/AERQBMIHwXA/sddefault.webp" alt="Regression with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn to perform regression analysis using Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Preparation&lt;/strong&gt;: Understanding the cleaned dataset and necessary columns for analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enabling the Tool&lt;/strong&gt;: How to enable the Data Analysis Tool Pack in Excel.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Types of Regression&lt;/strong&gt;: Differences between simple and multiple linear regression.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Setting Up Regression&lt;/strong&gt;: Steps to input dependent (new deaths) and independent variables (new cases, new tests, new vaccinations, stringency index) for the analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interpreting Output&lt;/strong&gt;: Reading the regression output, focusing on adjusted R-squared, significance value (F-test), and P-values.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coefficient Interpretation&lt;/strong&gt;: Understanding the impact of each independent variable on the dependent variable, including scaling factors (per 1000 units).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model Evaluation&lt;/strong&gt;: Evaluating the model based on significance values and understanding the implications of unexpected results (e.g., stringency index).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Further Analysis&lt;/strong&gt;: Recognizing the need for additional analysis when encountering unexpected or inconclusive results.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/rest-apis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/rest-apis/</guid><description>&lt;h2 id="rest-apis"&gt;REST APIs&lt;a class="anchor" href="#rest-apis"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;REST (Representational State Transfer) APIs are the standard way to build web services that allow different systems to communicate over HTTP. They use standard HTTP methods and JSON for data exchange.&lt;/p&gt;
&lt;p&gt;Watch this comprehensive introduction to REST APIs (52 min):&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/qbLc5a9jdXo"&gt;&lt;img src="https://i.ytimg.com/vi_webp/qbLc5a9jdXo/sddefault.webp" alt="REST API Crash Course - Introduction &amp;#43; Full Python API Tutorial (52)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Key Concepts:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;HTTP Methods&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;GET&lt;/code&gt;: Retrieve data&lt;/li&gt;
&lt;li&gt;&lt;code&gt;POST&lt;/code&gt;: Create new data&lt;/li&gt;
&lt;li&gt;&lt;code&gt;PUT/PATCH&lt;/code&gt;: Update existing data&lt;/li&gt;
&lt;li&gt;&lt;code&gt;DELETE&lt;/code&gt;: Remove data&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Status Codes&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;2xx&lt;/code&gt;: Success (200 OK, 201 Created)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;4xx&lt;/code&gt;: Client errors (400 Bad Request, 404 Not Found)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;5xx&lt;/code&gt;: Server errors (500 Internal Server Error)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal REST API using FastAPI. Run this &lt;code&gt;server.py&lt;/code&gt; script via &lt;code&gt;uv run server.py&lt;/code&gt;:&lt;/p&gt;</description></item><item><title/><link>/2026-05/retrieval-augmented-generation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/retrieval-augmented-generation/</guid><description>&lt;h2 id="retrieval-augmented-generation"&gt;Retrieval Augmented Generation&lt;a class="anchor" href="#retrieval-augmented-generation"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/T-D1OfcDW1M"&gt;&lt;img src="https://i.ytimg.com/vi_webp/T-D1OfcDW1M/sddefault.webp" alt="What is Retrieval-Augmented Generation (RAG)? (6 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Watch the walkthrough above, then dive into the notebook for step-by-step practice.&lt;/p&gt;
&lt;p&gt;You will learn to implement Retrieval Augmented Generation (RAG) to enhance language models&amp;rsquo; responses by incorporating relevant context, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LLM Context Limitations&lt;/strong&gt;: Understanding the constraints of context windows in large language models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Retrieval Augmented Generation&lt;/strong&gt;: The technique of retrieving and using relevant documents to enhance model responses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embeddings&lt;/strong&gt;: How to convert text into numerical representations that are used for similarity calculations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Similarity Search&lt;/strong&gt;: Finding the most relevant documents by calculating cosine similarity between embeddings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI API Integration&lt;/strong&gt;: Using the OpenAI API to generate responses based on the most relevant documents.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tourist Recommendation Bot&lt;/strong&gt;: Building a bot that recommends tourist attractions based on user interests using embeddings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Next Steps for Implementation&lt;/strong&gt;: Insights into scaling the solution with a vector database, re-rankers, and improved prompts for better accuracy and efficiency.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/revealjs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/revealjs/</guid><description>&lt;h2 id="revealjs-web-based-presentations"&gt;RevealJS: Web-Based Presentations&lt;a class="anchor" href="#revealjs-web-based-presentations"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=a6ioNtv2H-E&amp;amp;list=PLoqZcxvpWzzf_C8QgHC9XbA-bn18qJ-QV"&gt;&lt;img src="https://i.ytimg.com/vi/a6ioNtv2H-E/sddefault.jpg" alt="Create Beautiful HTML Presentations with RevealJS (25 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://revealjs.com/"&gt;RevealJS&lt;/a&gt; is a powerful HTML presentation framework that turns web pages into interactive slideshows. It&amp;rsquo;s particularly useful for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Presentations&lt;/strong&gt;: Interactive charts, live data demos&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Technical Talks&lt;/strong&gt;: Code with syntax highlighting, math equations&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Web Portfolios&lt;/strong&gt;: Showcase projects with live demos&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Educational Content&lt;/strong&gt;: Interactive learning materials&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Conference Talks&lt;/strong&gt;: Remote-friendly with speaker notes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Documentation&lt;/strong&gt;: API and library showcases&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This tutorial covers creating interactive presentations with RevealJS.&lt;/p&gt;</description></item><item><title/><link>/2026-05/scheduled-scraping-with-github-actions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scheduled-scraping-with-github-actions/</guid><description>&lt;h2 id="scheduled-scraping-with-github-actions"&gt;Scheduled Scraping with GitHub Actions&lt;a class="anchor" href="#scheduled-scraping-with-github-actions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;GitHub Actions provides an excellent platform for running web scrapers on a schedule. This tutorial shows how to automate data collection from websites using GitHub Actions workflows.&lt;/p&gt;
&lt;h3 id="key-concepts"&gt;Key Concepts&lt;a class="anchor" href="#key-concepts"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scheduling&lt;/strong&gt;: Use &lt;a href="https://docs.github.com/en/actions/using-workflows/events-that-trigger-workflows#schedule"&gt;cron syntax&lt;/a&gt; to run scrapers at specific times&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dependencies&lt;/strong&gt;: Install required packages like &lt;code&gt;httpx&lt;/code&gt;, &lt;code&gt;lxml&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Storage&lt;/strong&gt;: Save scraped data to files and commit back to the repository&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Error Handling&lt;/strong&gt;: Implement robust error handling for network issues and HTML parsing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate Limiting&lt;/strong&gt;: Respect website terms of service and implement delays between requests&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&amp;rsquo;s a sample &lt;code&gt;scrape.py&lt;/code&gt; that scrapes the IMDb Top 250 movies using httpx and lxml:&lt;/p&gt;</description></item><item><title/><link>/2026-05/scraping-emarketer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-emarketer/</guid><description>&lt;h2 id="scraping-emarketer"&gt;Scraping emarketer&lt;a class="anchor" href="#scraping-emarketer"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In this live scraping session, we explore a real-life scenario where Straive had to scrape data from emarketer.com for a demo. This is a fairly realistic and representative way of how one might go about scraping a website.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/ZzUsDE1XjhE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/ZzUsDE1XjhE/sddefault.webp" alt="Live scraping session" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scraping&lt;/strong&gt;: How to extract data from web pages, including constructing URLs, fetching page content, and parsing HTML using packages like &lt;a href="https://lxml.de/"&gt;&lt;code&gt;lxml&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://www.python-httpx.org/"&gt;&lt;code&gt;httpx&lt;/code&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Caching&lt;/strong&gt;: Implementing a caching strategy to avoid redundant data fetching for efficiency and reliability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Error Handling and Debugging&lt;/strong&gt;: Practical tips for troubleshooting, such as using liberal print statements, breakpoints for in-depth debugging, and the concept of &amp;ldquo;rubber duck debugging&amp;rdquo; to clarify problems.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLMs&lt;/strong&gt;: Benefits of Gemini / ChatGPT for code suggestions and troubleshooting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real-World Application&lt;/strong&gt;: How quick proofs of concept to showcase capabilities to clients, emphasizing practice over theory.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/scraping-imdb-with-javascript/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-imdb-with-javascript/</guid><description>&lt;h2 id="scraping-imdb-with-javascript"&gt;Scraping IMDb with JavaScript&lt;a class="anchor" href="#scraping-imdb-with-javascript"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/YVIKZqZIcCo"&gt;&lt;img src="https://i.ytimg.com/vi_webp/YVIKZqZIcCo/sddefault.webp" alt="Scraping the IMDb with Browser JavaScript" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to scrape the &lt;a href="https://www.imdb.com/chart/top"&gt;IMDb Top 250 movies&lt;/a&gt; directly in the browser using JavaScript on the Chrome DevTools, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Access Developer Tools&lt;/strong&gt;: Use F12 or right-click &amp;gt; Inspect to open developer tools in Chrome or Edge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inspect Elements&lt;/strong&gt;: Identify and inspect HTML elements using the Elements tab.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Query Selectors&lt;/strong&gt;: Use &lt;code&gt;document.querySelectorAll&lt;/code&gt; and &lt;code&gt;document.querySelector&lt;/code&gt; to find elements by CSS class.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extract Text Content&lt;/strong&gt;: Retrieve text content from elements using JavaScript.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Functional Programming&lt;/strong&gt;: Apply &lt;a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array/map"&gt;map&lt;/a&gt;
and &lt;a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Functions/Arrow_functions"&gt;arrow functions&lt;/a&gt;
for concise data processing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Structuring&lt;/strong&gt;: Collect and format data into an array of arrays.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Copying Data&lt;/strong&gt;: Use the copy function to transfer data to the clipboard.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Convert to Spreadsheet&lt;/strong&gt;: Use online tools to convert JSON data to CSV or Excel format.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text Manipulation&lt;/strong&gt;: Perform text splitting and cleaning in Excel for final data formatting.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links and references:&lt;/p&gt;</description></item><item><title/><link>/2026-05/scraping-live-sessions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-live-sessions/</guid><description>&lt;h2 id="scraping-live-sessions"&gt;Scraping: Live Sessions&lt;a class="anchor" href="#scraping-live-sessions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/cAriusuJsmw"&gt;&lt;img src="https://i.ytimg.com/vi_webp/cAriusuJsmw/sddefault.webp" alt="Intro to Web scraping and HTML" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Fundamentals of web scraping with urllib and BeautifulSoup&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/I3auyTYORTs"&gt;&lt;img src="https://i.ytimg.com/vi_webp/I3auyTYORTs/sddefault.webp" alt="Fundamentals of web scraping with urllib and BeautifulSoup" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Intermediate web scraping use of cookies&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/DryMIxMf3VU"&gt;&lt;img src="https://i.ytimg.com/vi_webp/DryMIxMf3VU/sddefault.webp" alt="Intermediate web scraping use of cookies" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;XML intro and scraping&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/8S_jvsjtaYg"&gt;&lt;img src="https://i.ytimg.com/vi_webp/8S_jvsjtaYg/sddefault.webp" alt="XML intro and scraping" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/scraping-pdfs-with-tabula/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-pdfs-with-tabula/</guid><description>&lt;h2 id="scraping-pdfs-with-tabula"&gt;Scraping PDFs with Tabula&lt;a class="anchor" href="#scraping-pdfs-with-tabula"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/yDoKlKyxClQ"&gt;&lt;img src="https://i.ytimg.com/vi_webp/yDoKlKyxClQ/sddefault.webp" alt="Scrape PDFs with Tabula Python library" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to scrape tables from PDFs using the &lt;code&gt;tabula&lt;/code&gt; Python library, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Import Libraries&lt;/strong&gt;: Use Beautiful Soup for URL parsing and Tabula for extracting tables from PDFs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Specify Save Location&lt;/strong&gt;: Mount Google Drive to save scraped PDFs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Identify PDF URLs&lt;/strong&gt;: Parse the given URL to identify and select all PDF links.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Download PDFs&lt;/strong&gt;: Loop through identified links, saving each PDF to the specified location.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extract Tables&lt;/strong&gt;: Use Tabula to read tabular content from the downloaded PDFs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Control Extraction Area&lt;/strong&gt;: Specify page and area coordinates to accurately extract tables, avoiding extraneous text.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Save Extracted Data&lt;/strong&gt;: Convert the extracted table data into structured formats like CSV for further analysis.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links and references:&lt;/p&gt;</description></item><item><title/><link>/2026-05/scraping-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-with-excel/</guid><description>&lt;h2 id="scraping-with-excel"&gt;Scraping with Excel&lt;a class="anchor" href="#scraping-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/OCl6UdpmzRQ"&gt;&lt;img src="https://i.ytimg.com/vi_webp/OCl6UdpmzRQ/sddefault.webp" alt="Weather Scraping with Excel: Get the Data" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to &lt;a href="https://support.microsoft.com/en-au/office/import-data-from-the-web-b13eed81-33fe-410d-9247-1747269c28e4"&gt;import tables on the web using Excel&lt;/a&gt;, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Import from Web&lt;/strong&gt;: Use the query feature in Excel to scrape data from websites.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Establishing Web Connections&lt;/strong&gt;: Connect Excel to a web page using a URL.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Using Query Editor&lt;/strong&gt;: Navigate the query editor to view and manage web data tables.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Loading Data&lt;/strong&gt;: Load data from the web into Excel for further manipulation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Transformation&lt;/strong&gt;: Remove unnecessary columns and transform data as needed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Applying Transformations&lt;/strong&gt;: Track applied steps in the sequence for reproducibility.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Refreshing Data&lt;/strong&gt;: Refresh the imported data to get the latest updates from the web.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/scraping-with-google-sheets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/scraping-with-google-sheets/</guid><description>&lt;h2 id="scraping-with-google-sheets"&gt;Scraping with Google Sheets&lt;a class="anchor" href="#scraping-with-google-sheets"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/eYQEk7XJM7s"&gt;&lt;img src="https://i.ytimg.com/vi_webp/eYQEk7XJM7s/sddefault.webp" alt="Scraping with Google Sheets" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to &lt;a href="https://support.google.com/docs/answer/3093339?hl=en"&gt;import tables on the web using Google Sheets&amp;rsquo;s &lt;code&gt;=IMPORTHTML()&lt;/code&gt; formula&lt;/a&gt;, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Import HTML Formula&lt;/strong&gt;: Use =IMPORTHTML(URL, &amp;ldquo;query&amp;rdquo;, index) to fetch tables or lists from a web page.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Granting Access&lt;/strong&gt;: Allow access for formulas to fetch data from external sources.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Checking Imported Data&lt;/strong&gt;: Verify if the imported table matches the data on the web page.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Errors&lt;/strong&gt;: Understand common issues and how to resolve them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sorting Data&lt;/strong&gt;: Copy imported data as values and sort it within Google Sheets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Freezing Rows&lt;/strong&gt;: Use frozen rows to maintain headers while sorting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Live Formulas&lt;/strong&gt;: Learn how web data updates automatically when the source changes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Other Import Functions&lt;/strong&gt;: IMPORTXML, IMPORTFEED, IMPORTRANGE, and IMPORTDATA for advanced data fetching options.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/splitting-text-in-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/splitting-text-in-excel/</guid><description>&lt;h2 id="splitting-text-in-excel"&gt;Splitting Text in Excel&lt;a class="anchor" href="#splitting-text-in-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/fQeADnqiOAg"&gt;&lt;img src="https://i.ytimg.com/vi_webp/fQeADnqiOAg/sddefault.webp" alt="Convert text-to-columns in Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to transform a single-column data set into multiple, organized columns based on specific delimiters using the &amp;ldquo;Text to Columns&amp;rdquo; feature.&lt;/p&gt;
&lt;p&gt;Here are links used in the video:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.senate.gov/legislative/votes_new.htm"&gt;US Senate Legislation - Votes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/spreadsheets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/spreadsheets/</guid><description>&lt;h2 id="spreadsheet-excel-google-sheets"&gt;Spreadsheet: Excel, Google Sheets&lt;a class="anchor" href="#spreadsheet-excel-google-sheets"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You&amp;rsquo;ll use spreadsheets for data cleaning and exploration. The most popular spreadsheet program is &lt;a href="https://www.microsoft.com/en-us/microsoft-365/excel"&gt;Microsoft Excel&lt;/a&gt; followed by &lt;a href="https://www.google.com/sheets/about/"&gt;Google Sheets&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;You may be already familiar with these. If not, make sure to learn the basics of both.&lt;/p&gt;
&lt;p&gt;Go through the &lt;a href="https://support.microsoft.com/en-us/office/excel-video-training-9bc05390-e94c-46af-a5b3-d7c22f6990bb"&gt;&lt;strong&gt;Microsoft Excel&lt;/strong&gt; video training&lt;/a&gt; and make sure you cover:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/create-a-new-workbook-ae99f19b-cecb-4aa0-92c8-7126d6212a83"&gt;Intro to Excel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/insert-or-delete-rows-and-columns-6f40e6e4-85af-45e0-b39d-65dd504a3246"&gt;Rows &amp;amp; columns&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/move-or-copy-cells-and-cell-contents-803d65eb-6a3e-4534-8c6f-ff12d1c4139e"&gt;Cells&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/available-number-formats-in-excel-0afe8f52-97db-41f1-b972-4b46e9f1e8d2"&gt;Formatting&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/overview-of-formulas-in-excel-ecfdc708-9162-49e8-b993-c311f47ca173"&gt;Formulas &amp;amp; Functions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/create-and-format-tables-e81aa349-b006-4f8a-9806-5af9df0ac664"&gt;Tables&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.microsoft.com/en-us/office/create-a-pivottable-to-analyze-worksheet-data-a9a84538-bfe9-40a9-a8e9-f99134456576"&gt;PivotTables&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Watch this video for an introduction to &lt;strong&gt;Google Sheets&lt;/strong&gt; (49 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/sqlite/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/sqlite/</guid><description>&lt;h2 id="database-sqlite"&gt;Database: SQLite&lt;a class="anchor" href="#database-sqlite"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Relational databases are used to store data in a structured way. You&amp;rsquo;ll often access databases created by others for analysis.&lt;/p&gt;
&lt;p&gt;PostgreSQL, MySQL, MS SQL, Oracle, etc. are popular databases. But the most installed database is &lt;a href="https://www.sqlite.org/index.html"&gt;SQLite&lt;/a&gt;. It&amp;rsquo;s embedded into many devices and apps (e.g. your phone, browser, etc.). It&amp;rsquo;s lightweight but very scalable and powerful.&lt;/p&gt;
&lt;p&gt;Watch these introductory videos to understand SQLite and how it&amp;rsquo;s used in Python (34 min):&lt;/p&gt;</description></item><item><title/><link>/2026-05/system-requirements/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/system-requirements/</guid><description>&lt;h1 id="system-requirements"&gt;System Requirements&lt;a class="anchor" href="#system-requirements"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This course requires the following:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="/2026-05/system-requirements/#software"&gt;Software&lt;/a&gt; to be installed on your computer.&lt;/li&gt;
&lt;li&gt;&lt;a href="/2026-05/system-requirements/#system-permissions"&gt;System permissions&lt;/a&gt; to be granted.&lt;/li&gt;
&lt;li&gt;Access to the &lt;a href="/2026-05/system-requirements/#websites"&gt;Websites&lt;/a&gt; listed below.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="software"&gt;Software&lt;a class="anchor" href="#software"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Core tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/"&gt;Visual Studio Code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.astral.sh/uv/getting-started/installation/"&gt;uv&lt;/a&gt; for Python&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nodejs.org/"&gt;NodeJS&lt;/a&gt; for JavaScript&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.docker.com/"&gt;Docker&lt;/a&gt; or &lt;a href="https://podman.io/getting-started/installation"&gt;Podman&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/features/copilot"&gt;GitHub Copilot&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Module-specific tools:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dbeaver.io/"&gt;DBeaver&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://duckdb.org/"&gt;DuckDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.microsoft.com/en-in/microsoft-365/excel"&gt;Excel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ffmpeg.org/"&gt;FFMpeg&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cli.github.com/"&gt;GitHub CLI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://desktop.github.com/"&gt;GitHub Desktop&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://llm.datasette.io/"&gt;llm&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://lynx.invisible-island.net/"&gt;Lynx&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://mupdf.com/"&gt;MuPDF&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://obsproject.com/"&gt;OBS Studio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ollama.com/"&gt;Ollama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openrefine.org"&gt;OpenRefine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pandoc.org/"&gt;Pandoc&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.xpdfreader.com/pdftotext-man.html"&gt;pdftotext&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://playwright.dev/python/docs/intro"&gt;Playwright&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.qgis.org/en/site/"&gt;QGIS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sqlite-utils.datasette.io/"&gt;sqlite-utils&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sqlitestudio.pl/"&gt;SQLiteStudio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://typesense.org/docs/guide/install-typesense.html"&gt;TypeSense&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://w3m.sourceforge.net/"&gt;w3m&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.gnu.org/software/wget/"&gt;wget&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Purfview/whisper-standalone-win/releases"&gt;Whisper&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp"&gt;yt-dlp&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="system-permissions"&gt;System permissions&lt;a class="anchor" href="#system-permissions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Open Developer Tools on Chrome / Edge - to use F12 or Inspect in browser&lt;/li&gt;
&lt;li&gt;Open port 8000-9000 for local development server&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pip install&lt;/code&gt; and &lt;code&gt;npm install&lt;/code&gt; to install packages&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="websites"&gt;Websites&lt;a class="anchor" href="#websites"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;!-- Generated via

grep -ohrE 'https?://[^/ )":]+' *.md | sed 's#https\?://##' | sed 's/^www\.//' | sort | uniq

--&gt;
&lt;ul&gt;
&lt;li&gt;accounts.google.com&lt;/li&gt;
&lt;li&gt;agents.md&lt;/li&gt;
&lt;li&gt;ai.google.dev&lt;/li&gt;
&lt;li&gt;aipipe.org&lt;/li&gt;
&lt;li&gt;aiproxy.sanand.workers.dev&lt;/li&gt;
&lt;li&gt;aircloak.com&lt;/li&gt;
&lt;li&gt;airflow.apache.org&lt;/li&gt;
&lt;li&gt;aistudio.google.com&lt;/li&gt;
&lt;li&gt;anthropic.com&lt;/li&gt;
&lt;li&gt;api-atlas.nomic.ai&lt;/li&gt;
&lt;li&gt;api.example.com&lt;/li&gt;
&lt;li&gt;apify.com&lt;/li&gt;
&lt;li&gt;api.jina.ai&lt;/li&gt;
&lt;li&gt;api.openai.com&lt;/li&gt;
&lt;li&gt;api.open-notify.org&lt;/li&gt;
&lt;li&gt;app.example.com&lt;/li&gt;
&lt;li&gt;app.flourish.studio&lt;/li&gt;
&lt;li&gt;apple.com&lt;/li&gt;
&lt;li&gt;appsource.microsoft.com&lt;/li&gt;
&lt;li&gt;arxiv.org&lt;/li&gt;
&lt;li&gt;aspiegel.com&lt;/li&gt;
&lt;li&gt;atlas.nomic.ai&lt;/li&gt;
&lt;li&gt;azure.microsoft.com&lt;/li&gt;
&lt;li&gt;base64decode.org&lt;/li&gt;
&lt;li&gt;bbc.com&lt;/li&gt;
&lt;li&gt;beautiful-soup-4.readthedocs.io&lt;/li&gt;
&lt;li&gt;bing.com&lt;/li&gt;
&lt;li&gt;blog.gramener.com&lt;/li&gt;
&lt;li&gt;bolt.new&lt;/li&gt;
&lt;li&gt;calendar.google.com&lt;/li&gt;
&lt;li&gt;cb-prod.seek.study.iitm.ac.in&lt;/li&gt;
&lt;li&gt;cdn.jsdelivr.net&lt;/li&gt;
&lt;li&gt;cdn.openai.com&lt;/li&gt;
&lt;li&gt;chat.deepseek.com&lt;/li&gt;
&lt;li&gt;chatgpt.com&lt;/li&gt;
&lt;li&gt;chat.openai.com&lt;/li&gt;
&lt;li&gt;chat.qwen.ai&lt;/li&gt;
&lt;li&gt;claude.ai&lt;/li&gt;
&lt;li&gt;cli.github.com&lt;/li&gt;
&lt;li&gt;cline.bot&lt;/li&gt;
&lt;li&gt;cloud.google.com&lt;/li&gt;
&lt;li&gt;cmdlinetips.com&lt;/li&gt;
&lt;li&gt;codeium.com&lt;/li&gt;
&lt;li&gt;code.visualstudio.com&lt;/li&gt;
&lt;li&gt;coffee-reviews.prayashm.com&lt;/li&gt;
&lt;li&gt;colab.research.google.com&lt;/li&gt;
&lt;li&gt;console.cloud.google.com&lt;/li&gt;
&lt;li&gt;context7.com&lt;/li&gt;
&lt;li&gt;continue.dev&lt;/li&gt;
&lt;li&gt;cors-test.codehappy.devs&lt;/li&gt;
&lt;li&gt;cran.r-project.org&lt;/li&gt;
&lt;li&gt;cursor.com&lt;/li&gt;
&lt;li&gt;cyberscoop.com&lt;/li&gt;
&lt;li&gt;dashboard.ngrok.com&lt;/li&gt;
&lt;li&gt;dataforseo.com&lt;/li&gt;
&lt;li&gt;data.gov&lt;/li&gt;
&lt;li&gt;data.gov.in&lt;/li&gt;
&lt;li&gt;data.gov.ru&lt;/li&gt;
&lt;li&gt;datameet.org&lt;/li&gt;
&lt;li&gt;datasetsearch.research.google.com&lt;/li&gt;
&lt;li&gt;datasets.imdbws.com&lt;/li&gt;
&lt;li&gt;datasette.io&lt;/li&gt;
&lt;li&gt;dbeaver.io&lt;/li&gt;
&lt;li&gt;dbfopener.com&lt;/li&gt;
&lt;li&gt;deepwiki.com&lt;/li&gt;
&lt;li&gt;desktop.github.com&lt;/li&gt;
&lt;li&gt;devanshikat.github.io&lt;/li&gt;
&lt;li&gt;developer.chrome.com&lt;/li&gt;
&lt;li&gt;developer.imdb.com&lt;/li&gt;
&lt;li&gt;developer.mozilla.org&lt;/li&gt;
&lt;li&gt;developers.googleblog.com&lt;/li&gt;
&lt;li&gt;developers.google.com&lt;/li&gt;
&lt;li&gt;dev.mysql.com&lt;/li&gt;
&lt;li&gt;discourse.onlinedegree.iitm.ac.in&lt;/li&gt;
&lt;li&gt;diva-gis.org&lt;/li&gt;
&lt;li&gt;dl.acm.org&lt;/li&gt;
&lt;li&gt;docker.com&lt;/li&gt;
&lt;li&gt;docs.anthropic.com&lt;/li&gt;
&lt;li&gt;docs.astral.sh&lt;/li&gt;
&lt;li&gt;docs.cursor.com&lt;/li&gt;
&lt;li&gt;docs.docker.com&lt;/li&gt;
&lt;li&gt;docs.getdbt.com&lt;/li&gt;
&lt;li&gt;docs.github.com&lt;/li&gt;
&lt;li&gt;docs.google.com&lt;/li&gt;
&lt;li&gt;docs.klarna.com&lt;/li&gt;
&lt;li&gt;docs.kumu.io&lt;/li&gt;
&lt;li&gt;docs.lovable.dev&lt;/li&gt;
&lt;li&gt;docs.nomic.ai&lt;/li&gt;
&lt;li&gt;docs.npmjs.com&lt;/li&gt;
&lt;li&gt;docs.python.org&lt;/li&gt;
&lt;li&gt;docs.python-requests.org&lt;/li&gt;
&lt;li&gt;docs.replit.com&lt;/li&gt;
&lt;li&gt;docs.sqlalchemy.org&lt;/li&gt;
&lt;li&gt;docs.together.ai&lt;/li&gt;
&lt;li&gt;docs.windsurf.com&lt;/li&gt;
&lt;li&gt;dopiaza.org&lt;/li&gt;
&lt;li&gt;drive.google.com&lt;/li&gt;
&lt;li&gt;drive.usercontent.google.com&lt;/li&gt;
&lt;li&gt;duckdb.org&lt;/li&gt;
&lt;li&gt;elastic.co&lt;/li&gt;
&lt;li&gt;en.wikipedia.org&lt;/li&gt;
&lt;li&gt;evanhahn.github.io&lt;/li&gt;
&lt;li&gt;example.com&lt;/li&gt;
&lt;li&gt;exam.sanand.workers.dev&lt;/li&gt;
&lt;li&gt;explainxkcd.com&lt;/li&gt;
&lt;li&gt;fastapi.tiangolo.com&lt;/li&gt;
&lt;li&gt;fastapi-users.github.io&lt;/li&gt;
&lt;li&gt;ffmpeg.lav.io&lt;/li&gt;
&lt;li&gt;ffmpeg.org&lt;/li&gt;
&lt;li&gt;flourish.studio&lt;/li&gt;
&lt;li&gt;flukeout.github.io&lt;/li&gt;
&lt;li&gt;fmwconcepts.com&lt;/li&gt;
&lt;li&gt;forms.gle&lt;/li&gt;
&lt;li&gt;freecodecamp.org&lt;/li&gt;
&lt;li&gt;gemini.google&lt;/li&gt;
&lt;li&gt;gemini.google.com&lt;/li&gt;
&lt;li&gt;generativelanguage.googleapis.com&lt;/li&gt;
&lt;li&gt;geopy.readthedocs.io&lt;/li&gt;
&lt;li&gt;getdbt.com&lt;/li&gt;
&lt;li&gt;github.com&lt;/li&gt;
&lt;li&gt;github.github.com&lt;/li&gt;
&lt;li&gt;gitingest.com&lt;/li&gt;
&lt;li&gt;gitlab.com&lt;/li&gt;
&lt;li&gt;gitlens.amod.io&lt;/li&gt;
&lt;li&gt;git-lfs.github.com&lt;/li&gt;
&lt;li&gt;git-scm.com&lt;/li&gt;
&lt;li&gt;gnu.org&lt;/li&gt;
&lt;li&gt;googlecloudcommunity.com&lt;/li&gt;
&lt;li&gt;google.com&lt;/li&gt;
&lt;li&gt;gramener.com&lt;/li&gt;
&lt;li&gt;groups.google.com&lt;/li&gt;
&lt;li&gt;heroku.com&lt;/li&gt;
&lt;li&gt;howstat.com&lt;/li&gt;
&lt;li&gt;httpbin.org&lt;/li&gt;
&lt;li&gt;httpie.io&lt;/li&gt;
&lt;li&gt;httrack.com&lt;/li&gt;
&lt;li&gt;hub.docker.com&lt;/li&gt;
&lt;li&gt;hub.getdbt.com&lt;/li&gt;
&lt;li&gt;huggingface.co&lt;/li&gt;
&lt;li&gt;imagemagick.org&lt;/li&gt;
&lt;li&gt;imageoptim.com&lt;/li&gt;
&lt;li&gt;imdb.com&lt;/li&gt;
&lt;li&gt;imgs.xkcd.com&lt;/li&gt;
&lt;li&gt;img.youtube.com&lt;/li&gt;
&lt;li&gt;i.ytimg.com&lt;/li&gt;
&lt;li&gt;jamendo.com&lt;/li&gt;
&lt;li&gt;jeroenjanssens.com&lt;/li&gt;
&lt;li&gt;jina.ai&lt;/li&gt;
&lt;li&gt;jmespath.org&lt;/li&gt;
&lt;li&gt;jqlang.org&lt;/li&gt;
&lt;li&gt;jqplay.org&lt;/li&gt;
&lt;li&gt;jsoneditoronline.org&lt;/li&gt;
&lt;li&gt;json-generator.com&lt;/li&gt;
&lt;li&gt;jsonlines.org&lt;/li&gt;
&lt;li&gt;jsonlint.com&lt;/li&gt;
&lt;li&gt;jsonpath.com&lt;/li&gt;
&lt;li&gt;jsonpathfinder.com&lt;/li&gt;
&lt;li&gt;json-schema.org&lt;/li&gt;
&lt;li&gt;jsonschemavalidator.net&lt;/li&gt;
&lt;li&gt;judgments.ecourts.gov.in&lt;/li&gt;
&lt;li&gt;kaggle.com&lt;/li&gt;
&lt;li&gt;keras.io&lt;/li&gt;
&lt;li&gt;khanacademy.org&lt;/li&gt;
&lt;li&gt;kumu.io&lt;/li&gt;
&lt;li&gt;learn.microsoft.com&lt;/li&gt;
&lt;li&gt;linkedin.com&lt;/li&gt;
&lt;li&gt;llm.datasette.io&lt;/li&gt;
&lt;li&gt;llmstxt.org&lt;/li&gt;
&lt;li&gt;localhost&lt;/li&gt;
&lt;li&gt;locator-service.api.bbci.co.uk&lt;/li&gt;
&lt;li&gt;lovable.dev&lt;/li&gt;
&lt;li&gt;lxml.de&lt;/li&gt;
&lt;li&gt;lynx.invisible-island.net&lt;/li&gt;
&lt;li&gt;macstories.net&lt;/li&gt;
&lt;li&gt;magickstudio.imagemagick.org&lt;/li&gt;
&lt;li&gt;mail.google.com&lt;/li&gt;
&lt;li&gt;makersuite.google.com&lt;/li&gt;
&lt;li&gt;mapshaper.org&lt;/li&gt;
&lt;li&gt;marimo.app&lt;/li&gt;
&lt;li&gt;marimo.io&lt;/li&gt;
&lt;li&gt;marketplace.visualstudio.com&lt;/li&gt;
&lt;li&gt;marp.app&lt;/li&gt;
&lt;li&gt;matplotlib.org&lt;/li&gt;
&lt;li&gt;medium.com&lt;/li&gt;
&lt;li&gt;microsoft.com&lt;/li&gt;
&lt;li&gt;mistral.ai&lt;/li&gt;
&lt;li&gt;modelcontextprotocol.io&lt;/li&gt;
&lt;li&gt;mongodb.com&lt;/li&gt;
&lt;li&gt;mupdf.com&lt;/li&gt;
&lt;li&gt;nchandrasekharr.github.io&lt;/li&gt;
&lt;li&gt;news.ycombinator.com&lt;/li&gt;
&lt;li&gt;ngrok.com&lt;/li&gt;
&lt;li&gt;nodejs.org&lt;/li&gt;
&lt;li&gt;nominatim.org&lt;/li&gt;
&lt;li&gt;notebooklm.google.com&lt;/li&gt;
&lt;li&gt;npmjs.com&lt;/li&gt;
&lt;li&gt;numpy.org&lt;/li&gt;
&lt;li&gt;observablehq.com&lt;/li&gt;
&lt;li&gt;obsproject.com&lt;/li&gt;
&lt;li&gt;ollama.com&lt;/li&gt;
&lt;li&gt;openai.com&lt;/li&gt;
&lt;li&gt;openpolicyagent.org&lt;/li&gt;
&lt;li&gt;openrefine.org&lt;/li&gt;
&lt;li&gt;openrouter.ai&lt;/li&gt;
&lt;li&gt;ourworldindata.org&lt;/li&gt;
&lt;li&gt;packaging.python.org&lt;/li&gt;
&lt;li&gt;pages.github.com&lt;/li&gt;
&lt;li&gt;pandas.pydata.org&lt;/li&gt;
&lt;li&gt;pandoc.org&lt;/li&gt;
&lt;li&gt;parquet.apache.org&lt;/li&gt;
&lt;li&gt;pillow.readthedocs.io&lt;/li&gt;
&lt;li&gt;platform.openai.com&lt;/li&gt;
&lt;li&gt;playwright.dev&lt;/li&gt;
&lt;li&gt;pngquant.org&lt;/li&gt;
&lt;li&gt;podman.io&lt;/li&gt;
&lt;li&gt;pokeapi.co&lt;/li&gt;
&lt;li&gt;posit.co&lt;/li&gt;
&lt;li&gt;postgresql.org&lt;/li&gt;
&lt;li&gt;postman.com&lt;/li&gt;
&lt;li&gt;pre-commit.com&lt;/li&gt;
&lt;li&gt;projector.tensorflow.org&lt;/li&gt;
&lt;li&gt;promptfoo.dev&lt;/li&gt;
&lt;li&gt;pydantic-docs.helpmanual.io&lt;/li&gt;
&lt;li&gt;pymotw.com&lt;/li&gt;
&lt;li&gt;pymupdf.readthedocs.io&lt;/li&gt;
&lt;li&gt;pypi.org&lt;/li&gt;
&lt;li&gt;python-httpx.org&lt;/li&gt;
&lt;li&gt;python-visualization.github.io&lt;/li&gt;
&lt;li&gt;qgis.org&lt;/li&gt;
&lt;li&gt;quotes.toscrape.com&lt;/li&gt;
&lt;li&gt;rasagy.in&lt;/li&gt;
&lt;li&gt;rattle.togaware.com&lt;/li&gt;
&lt;li&gt;raw.githubusercontent.com&lt;/li&gt;
&lt;li&gt;rawgraphs.io&lt;/li&gt;
&lt;li&gt;realpython.com&lt;/li&gt;
&lt;li&gt;redis.io&lt;/li&gt;
&lt;li&gt;relational-data.org&lt;/li&gt;
&lt;li&gt;revealjs.com&lt;/li&gt;
&lt;li&gt;rishabhmakes.github.io&lt;/li&gt;
&lt;li&gt;roocode.com&lt;/li&gt;
&lt;li&gt;sanand0.github.io&lt;/li&gt;
&lt;li&gt;s-anand.net&lt;/li&gt;
&lt;li&gt;scikit-learn.org&lt;/li&gt;
&lt;li&gt;scikit-network.readthedocs.io&lt;/li&gt;
&lt;li&gt;screentogif.com&lt;/li&gt;
&lt;li&gt;seaborn.pydata.org&lt;/li&gt;
&lt;li&gt;seek.onlinedegree.iitm.ac.in&lt;/li&gt;
&lt;li&gt;senate.gov&lt;/li&gt;
&lt;li&gt;shapechef.com&lt;/li&gt;
&lt;li&gt;sharp.pixelplumbing.com&lt;/li&gt;
&lt;li&gt;simonwillison.net&lt;/li&gt;
&lt;li&gt;soundcloud.com&lt;/li&gt;
&lt;li&gt;sourceforge.net&lt;/li&gt;
&lt;li&gt;sqlite.org&lt;/li&gt;
&lt;li&gt;sqlitestudio.pl&lt;/li&gt;
&lt;li&gt;sqlite-utils.datasette.io&lt;/li&gt;
&lt;li&gt;sqlmodel.tiangolo.com&lt;/li&gt;
&lt;li&gt;sqlzoo.net&lt;/li&gt;
&lt;li&gt;squoosh.app&lt;/li&gt;
&lt;li&gt;stackoverflow.blog&lt;/li&gt;
&lt;li&gt;stackoverflow.com&lt;/li&gt;
&lt;li&gt;statsmodels.org&lt;/li&gt;
&lt;li&gt;stats.stackexchange.com&lt;/li&gt;
&lt;li&gt;stedolan.github.io&lt;/li&gt;
&lt;li&gt;story-b0f1c.web.app&lt;/li&gt;
&lt;li&gt;storytellingwithdata.com&lt;/li&gt;
&lt;li&gt;study.iitm.ac.in&lt;/li&gt;
&lt;li&gt;superlinked.com&lt;/li&gt;
&lt;li&gt;support.anthropic.com&lt;/li&gt;
&lt;li&gt;support.apple.com&lt;/li&gt;
&lt;li&gt;support.google.com&lt;/li&gt;
&lt;li&gt;support.microsoft.com&lt;/li&gt;
&lt;li&gt;support.office.com&lt;/li&gt;
&lt;li&gt;survey.stackoverflow.co&lt;/li&gt;
&lt;li&gt;swagger.io&lt;/li&gt;
&lt;li&gt;tableau.com&lt;/li&gt;
&lt;li&gt;tabula-py.readthedocs.io&lt;/li&gt;
&lt;li&gt;tds.s-anand.net&lt;/li&gt;
&lt;li&gt;textblob.readthedocs.io&lt;/li&gt;
&lt;li&gt;texttospeech.googleapis.com&lt;/li&gt;
&lt;li&gt;thoughtworks.com&lt;/li&gt;
&lt;li&gt;timeanddate.com&lt;/li&gt;
&lt;li&gt;tools.s-anand.net&lt;/li&gt;
&lt;li&gt;tracebit.com&lt;/li&gt;
&lt;li&gt;twitter.com&lt;/li&gt;
&lt;li&gt;typesense.org&lt;/li&gt;
&lt;li&gt;unstructured.io&lt;/li&gt;
&lt;li&gt;upload.wikimedia.org&lt;/li&gt;
&lt;li&gt;url.com&lt;/li&gt;
&lt;li&gt;us-central1-aiplatform.googleapis.com&lt;/li&gt;
&lt;li&gt;vercel.com&lt;/li&gt;
&lt;li&gt;w3m.sourceforge.net&lt;/li&gt;
&lt;li&gt;w3.org&lt;/li&gt;
&lt;li&gt;weather-broker-cdn.api.bbci.co.uk&lt;/li&gt;
&lt;li&gt;web.archive.org&lt;/li&gt;
&lt;li&gt;webmaster.petalsearch.com&lt;/li&gt;
&lt;li&gt;whitehouse.gov&lt;/li&gt;
&lt;li&gt;whois.com&lt;/li&gt;
&lt;li&gt;wikipedia.readthedocs.io&lt;/li&gt;
&lt;li&gt;windsurf.com&lt;/li&gt;
&lt;li&gt;wordlift.io&lt;/li&gt;
&lt;li&gt;xpdfreader.com&lt;/li&gt;
&lt;li&gt;yoast.com&lt;/li&gt;
&lt;li&gt;your-app.vercel.app&lt;/li&gt;
&lt;li&gt;youtu.be&lt;/li&gt;
&lt;li&gt;youtube.com&lt;/li&gt;
&lt;li&gt;youtubetranscript.com&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/tds-gpt-reviewer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/tds-gpt-reviewer/</guid><description>&lt;h1 id="tds-gpt-reviewer"&gt;TDS GPT Reviewer&lt;a class="anchor" href="#tds-gpt-reviewer"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;After the later parts of this course&amp;rsquo;s contents were written, we ran it through a &lt;a href="https://chatgpt.com/g/g-6777656ed3b8819187b6f17d9f343853-technical-content-reviewer"&gt;Technical Content Reviewer GPT&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Take a look at the GPT&amp;rsquo;s instructions. These were generated by the &lt;a href="https://platform.openai.com/docs/guides/prompt-generation"&gt;OpenAI Prompt Generation&lt;/a&gt; tool.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-markdown" data-lang="markdown"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;As a &lt;span style="font-weight:bold"&gt;**Content Reviewer**&lt;/span&gt; for a high school–level course on Tools in Data Science, your job is to evaluate provided content (such as text, code snippets, or references) with a focus on correctness, clarity, and conciseness, and offer actionable feedback for improvement.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;1.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Check for Correctness and Consistency**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Verify technical and factual accuracy.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Ensure internal consistency without contradictions.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;2.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Check for Clarity and Approachability**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Ensure content is understandable for a high school student with limited prior knowledge.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Identify and simplify jargon or advanced concepts.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;3.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Check for Conciseness**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Assess if content is direct and free of unnecessary verbosity.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Identify areas for streamlining to enhance readability.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;4.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Provide Feedback for Improvement**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Offer actionable suggestions for fixing, clarifying, or reorganizing content.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Propose alternative phrasing if text is vague, complex, or verbose.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;# Steps
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;1.&lt;/span&gt; Carefully read the entire content before forming conclusions.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;2.&lt;/span&gt; List factual inconsistencies or missing details causing confusion.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;3.&lt;/span&gt; Suggest simpler terms or analogies for complex language.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;4.&lt;/span&gt; Point out unnecessary repetition or filler text.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;5.&lt;/span&gt; Provide direct examples of how to improve the highlighted issues.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;# Output Format
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Respond using &lt;span style="font-weight:bold"&gt;**Markdown**&lt;/span&gt; with the following structure:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;1.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Summary of Findings**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; A concise paragraph outlining overall strengths and weaknesses.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;2.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Detailed Review**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; **Correctness and Consistency**: Note factual errors or inconsistencies, suggesting corrections.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; **Clarity and Approachability**: Identify overly advanced or unclear sections, offering simpler alternatives.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; **Conciseness**: Highlight long or repetitive sections with suggestions for tightening the text.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;3.&lt;/span&gt; &lt;span style="font-weight:bold"&gt;**Actionable Improvement Suggestions**&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Provide specific sentences, bullet points, or rewritten examples to illustrate improvements.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;# Notes
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;-&lt;/span&gt; Maintain a constructive review tone, not content generation.
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;- Even if content is perfect, confirm with suggestions for minor improvements (e.g., adding an example or clarifying a subtle point).&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="content-creation-prompts"&gt;Content creation prompts&lt;a class="anchor" href="#content-creation-prompts"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In addition, here are a few prompts used to create the content:&lt;/p&gt;</description></item><item><title/><link>/2026-05/tds-ta-instructions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/tds-ta-instructions/</guid><description>&lt;h1 id="tds-ta-instructions"&gt;TDS TA Instructions&lt;a class="anchor" href="#tds-ta-instructions"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;The TDS TA is a virtual assistant that helps you with your doubts.&lt;/p&gt;
&lt;p&gt;It has been trained on course content created as follows:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Clone the course repository&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git clone https://github.com/sanand0/tools-in-data-science-public.git
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd tools-in-data-science-public
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Create a prompt file for the TA&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;PYTHONUTF8&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt; uvx files-to-prompt --cxml *.md -o tds-content.xml
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Replace the source with the URL of the course&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sed -i &lt;span style="color:#e6db74"&gt;&amp;#34;s/&amp;lt;source&amp;gt;/&amp;lt;source&amp;gt;https:\/\/tds.s-anand.net\/#\//g&amp;#34;&lt;/span&gt; tds-content.xml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Additionally, we visit each of the evaluation links on &lt;a href="https://exam.sanand.workers.dev/"&gt;https://exam.sanand.workers.dev/&lt;/a&gt;, &lt;a href="https://tools.s-anand.net/page2md/"&gt;copy it as Markdown&lt;/a&gt;, and add it to the content, called &lt;code&gt;ga1.md&lt;/code&gt;, &lt;code&gt;ga2.md&lt;/code&gt;, etc.&lt;/p&gt;</description></item><item><title/><link>/2026-05/topic-modeling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/topic-modeling/</guid><description>&lt;h2 id="topic-modeling"&gt;Topic Modeling&lt;a class="anchor" href="#topic-modeling"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/eQUNhq91DlI"&gt;&lt;img src="https://i.ytimg.com/vi_webp/eQUNhq91DlI/sddefault.webp" alt="LLM Topic Modeling" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn to use text embeddings to find text similarity and use that to create topics automatically from text, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Embeddings&lt;/strong&gt;: How large language models convert text into numerical representations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Similarity Measurement&lt;/strong&gt;: Understanding how similar embeddings indicate similar meanings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embedding Visualization&lt;/strong&gt;: Using tools like Tensorflow Projector to visualize embedding spaces.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embedding Applications&lt;/strong&gt;: Using embeddings for tasks like classification and clustering.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI Embeddings&lt;/strong&gt;: Using OpenAI&amp;rsquo;s API to generate embeddings for text.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model Comparison&lt;/strong&gt;: Exploring different embedding models and their strengths and weaknesses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cosine Similarity&lt;/strong&gt;: Calculating cosine similarity between embeddings for more reliable similarity measures.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embedding Cost&lt;/strong&gt;: Understanding the cost of generating embeddings using OpenAI&amp;rsquo;s API.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embedding Range&lt;/strong&gt;: Understanding the range of values in embeddings and their significance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/transforming-images/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/transforming-images/</guid><description>&lt;h2 id="transforming-images"&gt;Transforming Images&lt;a class="anchor" href="#transforming-images"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;h3 id="image-processing-with-pil-pillow"&gt;Image Processing with PIL (Pillow)&lt;a class="anchor" href="#image-processing-with-pil-pillow"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://youtu.be/6Qs3wObeWwc"&gt;&lt;img src="https://i.ytimg.com/vi_webp/6Qs3wObeWwc/sddefault.webp" alt="Python Tutorial: Image Manipulation with Pillow (16 min)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://pypi.org/project/pillow/"&gt;Pillow&lt;/a&gt; is Python&amp;rsquo;s leading library for image processing, offering powerful tools for editing, analyzing, and generating images. It handles various formats (PNG, JPEG, GIF, etc.) and provides operations from basic resizing to complex filters.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a minimal example showing common operations:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# /// script&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# requires-python = &amp;#34;&amp;gt;=3.11&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# dependencies = [&amp;#34;Pillow&amp;#34;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# ///&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; PIL &lt;span style="color:#f92672"&gt;import&lt;/span&gt; Image, ImageEnhance, ImageFilter
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;async&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;process_image&lt;/span&gt;(path: str) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; Image&lt;span style="color:#f92672"&gt;.&lt;/span&gt;Image:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;Process an image with basic enhancements.&amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; Image&lt;span style="color:#f92672"&gt;.&lt;/span&gt;open(path) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; img:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Convert to RGB to ensure compatibility&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; img &lt;span style="color:#f92672"&gt;=&lt;/span&gt; img&lt;span style="color:#f92672"&gt;.&lt;/span&gt;convert(&lt;span style="color:#e6db74"&gt;&amp;#34;RGB&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Resize maintaining aspect ratio&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; img&lt;span style="color:#f92672"&gt;.&lt;/span&gt;thumbnail((&lt;span style="color:#ae81ff"&gt;800&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;800&lt;/span&gt;))
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Apply enhancements&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; img &lt;span style="color:#f92672"&gt;=&lt;/span&gt; ImageEnhance&lt;span style="color:#f92672"&gt;.&lt;/span&gt;Contrast(img)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;enhance(&lt;span style="color:#ae81ff"&gt;1.2&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; img&lt;span style="color:#f92672"&gt;.&lt;/span&gt;filter(ImageFilter&lt;span style="color:#f92672"&gt;.&lt;/span&gt;SHARPEN)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; __name__ &lt;span style="color:#f92672"&gt;==&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;__main__&amp;#34;&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#f92672"&gt;import&lt;/span&gt; asyncio
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; img &lt;span style="color:#f92672"&gt;=&lt;/span&gt; asyncio&lt;span style="color:#f92672"&gt;.&lt;/span&gt;run(process_image(&lt;span style="color:#e6db74"&gt;&amp;#34;input.jpg&amp;#34;&lt;/span&gt;))
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; img&lt;span style="color:#f92672"&gt;.&lt;/span&gt;save(&lt;span style="color:#e6db74"&gt;&amp;#34;output.jpg&amp;#34;&lt;/span&gt;, quality&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;85&lt;/span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Key features and techniques you&amp;rsquo;ll learn:&lt;/p&gt;</description></item><item><title/><link>/2026-05/unicode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/unicode/</guid><description>&lt;h2 id="unicode"&gt;Unicode&lt;a class="anchor" href="#unicode"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Ever noticed when you copy-paste some text and get garbage symbols? Or see garbage when you load a CSV file? This video explains why. It covers how computers store text (called character encoding) and why it sometimes goes wonky.&lt;/p&gt;
&lt;p&gt;Learn about ASCII (the original 7-bit encoding system that could only handle 128 characters), why that wasn&amp;rsquo;t enough for global languages, and how modern solutions like Unicode save the day by letting us use any character from any language.&lt;/p&gt;</description></item><item><title/><link>/2026-05/uv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/uv/</guid><description>&lt;h2 id="python-tools-uv"&gt;Python tools: uv&lt;a class="anchor" href="#python-tools-uv"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://docs.astral.sh/uv/getting-started/installation/"&gt;Install uv&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.astral.sh/uv/"&gt;&lt;code&gt;uv&lt;/code&gt;&lt;/a&gt; is a fast Python package and project manager that&amp;rsquo;s becoming the standard for running Python scripts. It replaces tools like pip, conda, pipx, poetry, pyenv, twine, and virtualenv into one, enabling:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Python Version Management&lt;/strong&gt;: uv installs and manages &lt;em&gt;multiple&lt;/em&gt; Python versions, allowing developers to specify and switch between versions seamlessly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Virtual Environment Handling&lt;/strong&gt;: It automates the creation and management of virtual environments, ensuring isolated and consistent development spaces for different projects.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dependency Management&lt;/strong&gt;: With support for the pyproject.toml format, uv enables precise specification of project dependencies. It maintains a universal lockfile, uv.lock, to ensure reproducible installations across different systems.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Project Execution&lt;/strong&gt;: The &lt;code&gt;uv run&lt;/code&gt; command allows for the execution of scripts and applications within the managed environment, streamlining development workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are some commonly used commands:&lt;/p&gt;</description></item><item><title/><link>/2026-05/vector-databases/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vector-databases/</guid><description>&lt;h2 id="vector-databases"&gt;Vector Databases&lt;a class="anchor" href="#vector-databases"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Vector databases are specialized databases that store and search vector embeddings efficiently.&lt;/p&gt;
&lt;p&gt;Use vector databases when your embeddings exceed available memory or when you want it run fast at scale. (This is important. If your code runs fast and fits in memory, you &lt;strong&gt;DON&amp;rsquo;T&lt;/strong&gt; need a vector database. You can can use &lt;code&gt;numpy&lt;/code&gt; for these tasks.)&lt;/p&gt;
&lt;p&gt;Vector databases are an evolving space.&lt;/p&gt;
&lt;p&gt;The first generation of vector databases were written in C and typically used an algorithm called &lt;a href="https://en.wikipedia.org/wiki/Hierarchical_navigable_small_world"&gt;HNSW&lt;/a&gt; (a way to approximately find the nearest neighbor). Some popular ones are:&lt;/p&gt;</description></item><item><title/><link>/2026-05/vercel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vercel/</guid><description>&lt;h2 id="serverless-hosting-vercel"&gt;Serverless hosting: Vercel&lt;a class="anchor" href="#serverless-hosting-vercel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;!--

Why Vercel? I evaluated from https://survey.stackoverflow.co/2024/technology#2-cloud-platforms

- AWS, Azure, Google Cloud are too complex for beginners
- Cloudflare (next most popular, widely admired) Python support is in beta
- Hetzner (most admired), Supabase (next most admired) do not have a serverless platform
- Fly.io (next most admired) does not have a free tier
- Heroku (used in previous terms) is the least admired
- Vercel is both popular, admired, growing, has a free plan, and a simple API

--&gt;
&lt;p&gt;Serverless platforms let you rent a single function instead of an entire machine. They&amp;rsquo;re perfect for small web tools that &lt;em&gt;don&amp;rsquo;t need to run all the time&lt;/em&gt;. Here are some common real-life uses:&lt;/p&gt;</description></item><item><title/><link>/2026-05/vibe-analysis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vibe-analysis/</guid><description>&lt;h1 id="vibe-analysis"&gt;Vibe Analysis&lt;a class="anchor" href="#vibe-analysis"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Vibe analysis is analyzing data as if the analysis itself doesn&amp;rsquo;t exist—you focus only on business outcomes, skipping intermediate steps.&lt;/p&gt;
&lt;p&gt;This approach emerged from &lt;strong&gt;vibe coding&lt;/strong&gt;, where you code as if code doesn&amp;rsquo;t exist. With vibe analysis, you give an LLM full context, state your goal, and review only the answers. The LLM handles exploration, cleaning, modeling, visualization, and deployment.&lt;/p&gt;
&lt;p&gt;Watch these comprehensive talks on vibe analysis and the changing role of data scientists:&lt;/p&gt;</description></item><item><title/><link>/2026-05/vibe-coding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vibe-coding/</guid><description>&lt;h1 id="vibe-coding"&gt;Vibe-coding&lt;a class="anchor" href="#vibe-coding"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;The rise of AI coding tools has introduced a new development philosophy: &lt;a href="https://en.wikipedia.org/wiki/Vibe_coding"&gt;&lt;strong&gt;vibe-coding&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Vibe coding is different from AI coding. It&amp;rsquo;s &amp;ldquo;fully giving in to the vibes&amp;hellip;&amp;rdquo; and &amp;ldquo;forgetting that the code even exists&amp;rdquo;, i.e., accept-all edits, don’t hand-edit diffs, etc. The goal is to get to a working prototype as fast as possible.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/MUEgLWcaXTE"&gt;&lt;img src="https://i.ytimg.com/vi_webp/MUEgLWcaXTE/sddefault.webp" alt="What is Vibe Coding?" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ship scrappy first drafts fast&lt;/strong&gt;: Focus on getting a working version quickly rather than polished code. AI can generate functional prototypes in minutes, allowing you to validate ideas before investing in refinement.&lt;/p&gt;</description></item><item><title/><link>/2026-05/vision-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vision-models/</guid><description>&lt;h2 id="vision-models"&gt;Vision Models&lt;a class="anchor" href="#vision-models"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/FgT_Mk_bakQ"&gt;&lt;img src="https://i.ytimg.com/vi_webp/FgT_Mk_bakQ/sddefault.webp" alt="LLM Vision Models" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to use LLMs to interpret images and extract useful information, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Setting Up Vision Models&lt;/strong&gt;: Integrate vision capabilities with LLMs using APIs like OpenAI&amp;rsquo;s Chat Completion.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sending Image URLs for Analysis&lt;/strong&gt;: Pass URLs or base64-encoded images to LLMs for processing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reading Image Responses&lt;/strong&gt;: Get detailed textual descriptions of images, from scenic landscapes to specific objects like cricketers or bank statements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracting Data from Images&lt;/strong&gt;: Convert extracted image data to various formats like Markdown tables or JSON arrays.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handling Model Hallucinations&lt;/strong&gt;: Address inaccuracies in extraction results, understanding how different prompts can affect output quality.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost Management for Vision Models&lt;/strong&gt;: Adjust detail settings (e.g., &amp;ldquo;detail: low&amp;rdquo;) to balance cost and output precision.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/visualizing-animated-data-with-flourish/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/visualizing-animated-data-with-flourish/</guid><description>&lt;h2 id="visualizing-animated-data-with-flourish"&gt;Visualizing Animated Data with Flourish&lt;a class="anchor" href="#visualizing-animated-data-with-flourish"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/JrnIu5Bm8i4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/JrnIu5Bm8i4/sddefault.webp" alt="Visualizing animated data with Flourish" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to create and customize animations in Flourish, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Core Animation Principles:&lt;/strong&gt; Understand how to use animation to both engage and inform, and learn the concept of &amp;ldquo;object constancy&amp;rdquo; to turn data visualizations into compelling data stories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Configuring Race Charts:&lt;/strong&gt; Master the animation controls for Line Chart Races and Bar Chart Races, including setting durations for timelines, transitions, and individual stages.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Morphing Between Chart Types:&lt;/strong&gt; Learn to use the Line Bar Pie template to create seamless morphing transitions between different chart types (e.g., line to streamgraph) within a Flourish story.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Animating Scatter Plots:&lt;/strong&gt; Discover how to set up and control animations in scatter plots, including using the &amp;ldquo;Animation stagger&amp;rdquo; setting to create dynamic, cascading dot movements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Controlling Animation Logic:&lt;/strong&gt; Understand how to manage animations for both related and unrelated datasets by toggling the &amp;ldquo;Only animate series with same name&amp;rdquo; option.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Animation in Other Templates:&lt;/strong&gt; Get an overview of the animation settings available in various other templates, including Survey, Heatmap, Spider, and Hierarchy charts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advanced Data Explorer Animations:&lt;/strong&gt; Learn about the powerful Data Explorer template (an Enterprise feature) and its ability to create advanced animations, such as morphing dots from a survey layout into their corresponding locations on a projection map.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/visualizing-animated-data-with-powerpoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/visualizing-animated-data-with-powerpoint/</guid><description>&lt;h2 id="visualizing-animated-data-with-powerpoint"&gt;Visualizing Animated Data with PowerPoint&lt;a class="anchor" href="#visualizing-animated-data-with-powerpoint"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/umHlPDFVWr0"&gt;&lt;img src="https://i.ytimg.com/vi_webp/umHlPDFVWr0/sddefault.webp" alt="Visualizing animated data with PowerPoint" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to create an animated bar chart race in PowerPoint, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Using the Morph Transition&lt;/strong&gt;: How to leverage PowerPoint&amp;rsquo;s Morph feature to animate the resizing and reordering of bars between slides, which is the core technique for the race effect.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Manual Bar Creation&lt;/strong&gt;: The process of using individual shapes (rectangles) instead of a standard chart and how to size them accurately to represent your data values.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adding and Syncing Audio&lt;/strong&gt;: How to record a voiceover for each slide, add background music that plays across the entire presentation, and sync them with the animations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automating Playback&lt;/strong&gt;: Setting up automatic slide transitions to create a self-playing, video-like presentation that advances after the narration for each slide is complete.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exporting as a Video&lt;/strong&gt;: The steps to save your final animation as an MPEG-4 video file directly from PowerPoint and the suggestion of alternative methods like OBS for higher quality.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prototyping Data Stories&lt;/strong&gt;: Understanding the manual effort required for this method and how it allows for careful control when prototyping and crafting a specific visual narrative.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="how-to-make-bar-chart-race-in-powerpoint"&gt;How to make bar chart race in PowerPoint&lt;a class="anchor" href="#how-to-make-bar-chart-race-in-powerpoint"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://blog.gramener.com/bar-chart-race-in-powerpoint/"&gt;Source: How to make a bar chart race in PowerPoint&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/visualizing-forecasts-with-excel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/visualizing-forecasts-with-excel/</guid><description>&lt;h2 id="visualizing-forecasts-with-excel"&gt;Visualizing Forecasts with Excel&lt;a class="anchor" href="#visualizing-forecasts-with-excel"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/judFpVgfsV4"&gt;&lt;img src="https://i.ytimg.com/vi_webp/judFpVgfsV4/sddefault.webp" alt="Visualizing forecasts with Excel" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.google.com/spreadsheets/d/1a6cSbmZKjX_ZzBsWWrPQwU_4KgRNMwc0/view#gid=1138079165"&gt;Excel File used in the video&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You&amp;rsquo;ll learn to perform exploratory data visualization on time-series financial data using Excel, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Visualizing Trends with Sparklines&lt;/strong&gt;: Use PivotTables and Sparklines to create compact, in-cell charts that reveal the performance trend of various financial securities (currencies, indices, commodities) over time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Forecasting and Measuring Volatility&lt;/strong&gt;: Learn to calculate key metrics like the average value, use the &lt;code&gt;GROWTH&lt;/code&gt; function to forecast the next day&amp;rsquo;s performance, and measure volatility using standard deviation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Analyzing Relationships with Scatter Plots&lt;/strong&gt;: Create scatter plots to visually investigate how two different securities move in relation to each other.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quantifying Correlation with R-Squared&lt;/strong&gt;: Add trendlines and display the R-squared value on a scatter plot to quantitatively measure how much of the movement in one variable is explained by another.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Using the Data Analysis ToolPak&lt;/strong&gt;: Learn how to enable and use Excel&amp;rsquo;s powerful Data Analysis ToolPak to perform more advanced statistical analysis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating a Correlation Matrix&lt;/strong&gt;: Use the ToolPak to efficiently generate a complete correlation matrix, showing the correlation coefficient for every pair of securities in the dataset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interpreting a Correlation Matrix with Conditional Formatting&lt;/strong&gt;: Apply color scales to the matrix to instantly identify strong positive correlations, negative correlations, and outliers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Discovering Insights and Forming Hypotheses&lt;/strong&gt;: Learn how to use these visualizations to spot patterns and form data-driven hypotheses, such as the high correlation between the currencies of neighboring countries.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&amp;rsquo;s another version of this analysis, delivered as part of a talk:&lt;/p&gt;</description></item><item><title/><link>/2026-05/visualizing-machine-learning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/visualizing-machine-learning/</guid><description>&lt;h2 id="visualizing-machine-learning"&gt;Visualizing Machine Learning&lt;a class="anchor" href="#visualizing-machine-learning"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/sORnCj52COw"&gt;&lt;img src="https://i.ytimg.com/vi_webp/sORnCj52COw/sddefault.webp" alt="Visualizing Machine Learning" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn about improving customer retention, understanding black box models, and using clustering for market segmentation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Churn Reduction&lt;/strong&gt;: Use decision trees to identify customers likely to leave.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost Efficiency&lt;/strong&gt;: Compare customer acquisition vs. retention costs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model Improvement&lt;/strong&gt;: Apply SVMs and neural networks for better accuracy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Project Challenges&lt;/strong&gt;: Understand issues with black box models in implementation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;K-Means Clustering&lt;/strong&gt;: Segment markets using demographic data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Visualization&lt;/strong&gt;: Interpret clustering results using maps and charts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Correlation Analysis&lt;/strong&gt;: Identify relationships between currency exchange rates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool Proficiency&lt;/strong&gt;: Utilize Excel, Python, and JavaScript for analysis and communication.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Practical Application&lt;/strong&gt;: Tailor marketing strategies based on cluster characteristics.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are the links used in the video:&lt;/p&gt;</description></item><item><title/><link>/2026-05/visualizing-network-data-with-kumu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/visualizing-network-data-with-kumu/</guid><description>&lt;h2 id="visualizing-network-data-with-kumu"&gt;Visualizing Network Data with Kumu&lt;a class="anchor" href="#visualizing-network-data-with-kumu"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/OndB17bigkc"&gt;&lt;img src="https://i.ytimg.com/vi_webp/OndB17bigkc/sddefault.webp" alt="Visualizing network data with Kumu" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://kumu.io"&gt;Kumu&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.imdb.com/non-commercial-datasets/"&gt;IMDb data&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://colab.research.google.com/drive/1CHR68fw7lZC9H2JtVW4LXpUvNwfM_VE-?usp=sharing"&gt;Jupyter Notebook&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://youtu.be/oi4fDzqsCes"&gt;&lt;img src="https://i.ytimg.com/vi_webp/oi4fDzqsCes/sddefault.webp" alt="Network analysis – filtering by year" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn about visualizing and analyzing relationships and networks using Kumu, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Understanding Kumu&lt;/strong&gt;: What Kumu is and its primary function as a tool to &lt;strong&gt;visualize complex relationships within data&lt;/strong&gt;. You&amp;rsquo;ll learn that it&amp;rsquo;s applicable beyond social analysis to &lt;strong&gt;any scenario involving relationships between entities&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Social Network Analysis&lt;/strong&gt;: How Kumu facilitates &lt;strong&gt;social network analysis&lt;/strong&gt;, helping to understand how different people and communities are connected, and identifying common interests.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Preparation for Kumu&lt;/strong&gt;: The process of preparing raw data, specifically &lt;strong&gt;IMDb actor data for Indian movies&lt;/strong&gt;, to be uploaded to Kumu. This includes filtering for only movies and Indian movies.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating Actor Networks&lt;/strong&gt;: How to construct an &lt;strong&gt;actor collaboration matrix&lt;/strong&gt; where each element denotes the number of movies two actors have acted in together. This involves a method using &lt;strong&gt;matrix multiplication&lt;/strong&gt; of a movie-actor matrix with its transpose.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimizing Sparse Matrices&lt;/strong&gt;: Understanding that actor collaboration matrices are often &lt;strong&gt;sparse (contain many zero entries)&lt;/strong&gt; and how to make computations fast and memory-efficient using the &lt;strong&gt;compressed sparse row (CSR) format&lt;/strong&gt; from the &lt;code&gt;scipy&lt;/code&gt; library in Python.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preparing Data for Kumu Upload&lt;/strong&gt;: How to convert the processed matrix data into the &amp;ldquo;from node to node&amp;rdquo; format, along with the &lt;strong&gt;strength of the connection&lt;/strong&gt; (number of shared movies), which is required for Kumu.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filtering Data Effectively&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Filtering by Year&lt;/strong&gt;: How to &lt;strong&gt;filter movie data by release year&lt;/strong&gt; (e.g., movies released after 1950) by converting the &amp;lsquo;start year&amp;rsquo; column to an integer data type, and troubleshooting common issues like newline characters (&lt;code&gt;/n&lt;/code&gt;) within string data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filtering by Language/Region&lt;/strong&gt;: How to filter for specific regions or languages, such as &lt;strong&gt;Indian movies&lt;/strong&gt;, by applying language options within data processing functions or by filtering a dedicated region data frame.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filtering Actor Pairs&lt;/strong&gt;: How to reduce the data size by filtering for actors who have acted in a minimum number of movies and for actor pairs who have acted together in a minimum number of movies.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visualizing and Analyzing Networks in Kumu&lt;/strong&gt;: How the prepared data creates a &lt;strong&gt;network of actors in Kumu&lt;/strong&gt;, allowing you to observe &lt;strong&gt;clusters&lt;/strong&gt; and understand direct and indirect connections.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exploring Network Connections&lt;/strong&gt;: How to &lt;strong&gt;search for specific actors&lt;/strong&gt; (e.g., Mohanlal) and examine their &lt;strong&gt;direct connections&lt;/strong&gt; (e.g., Mohanlal and Saikumar acted in 8 movies) and &lt;strong&gt;indirect connections&lt;/strong&gt; within the network.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Introduction to Community Detection&lt;/strong&gt;: A brief mention of &lt;strong&gt;community detection&lt;/strong&gt; as a method to identify groups within the network and understand their interconnections, to be explored in other tutorials.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;General Data Science Practices&lt;/strong&gt;: The importance of using resources like &lt;strong&gt;Google and documentation&lt;/strong&gt; for problem-solving, even for seemingly simple tasks, and the necessity of ensuring correct data types for operations.&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title/><link>/2026-05/vscode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/vscode/</guid><description>&lt;h2 id="editor-vs-code"&gt;Editor: VS Code&lt;a class="anchor" href="#editor-vs-code"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Your editor is the most important tool in your arsenal. That&amp;rsquo;s where you&amp;rsquo;ll spend most of your time. Make sure you&amp;rsquo;re comfortable with it.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://code.visualstudio.com/"&gt;&lt;strong&gt;Visual Studio Code&lt;/strong&gt;&lt;/a&gt; is, &lt;em&gt;by far&lt;/em&gt;, the most popular code editor today. According to the &lt;a href="https://survey.stackoverflow.co/2025/technology#1-dev-id-es"&gt;2025 StackOverflow Survey&lt;/a&gt; ~75% of developers use it. We recommend you learn it well. Even if you use another editor, you&amp;rsquo;ll be working with others who use it, and it&amp;rsquo;s a good idea to have some exposure.&lt;/p&gt;</description></item><item><title/><link>/2026-05/web-automation-with-playwright/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/web-automation-with-playwright/</guid><description>&lt;h2 id="web-scraping-with-playwright-in-python"&gt;Web Scraping with Playwright in Python&lt;a class="anchor" href="#web-scraping-with-playwright-in-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Scrape JavaScript‑heavy sites effortlessly with Playwright.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/biFzRHk4xpY"&gt;&lt;img src="https://i.ytimg.com/vi_webp/biFzRHk4xpY/sddefault.webp" alt="🤖 Playwright: Advanced Web Scraping in Python (14 min)" /&gt;&lt;/a&gt; (&lt;a href="https://www.youtube.com/watch?v=biFzRHk4xpY"&gt;youtube.com&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;Playwright offers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;JavaScript rendering&lt;/strong&gt;: Executes page scripts so you scrape only after content appears. (&lt;a href="https://playwright.dev/python/docs/intro"&gt;playwright.dev&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Headless &amp;amp; headed modes&lt;/strong&gt;: Run without UI or in a real browser for debugging. (&lt;a href="https://playwright.dev/python/docs/intro"&gt;playwright.dev&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auto‑waiting &amp;amp; retry&lt;/strong&gt;: Built‑in locators reduce flakiness. (&lt;a href="https://playwright.dev/python/docs/locators"&gt;playwright.dev&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi‑browser support&lt;/strong&gt;: Chromium, Firefox, WebKit—all from one API. (&lt;a href="https://playwright.dev/python/docs/intro"&gt;playwright.dev&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="example-scraping-a-jsrendered-site"&gt;Example: Scraping a JS‑Rendered Site&lt;a class="anchor" href="#example-scraping-a-jsrendered-site"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;p&gt;We’ll scrape &lt;a href="https://quotes.toscrape.com/js/"&gt;Quotes to Scrape (JS)&lt;/a&gt;—a site that loads quotes via JavaScript, so a simple &lt;code&gt;requests&lt;/code&gt; call gets only an empty shell (&lt;a href="https://quotes.toscrape.com/js/"&gt;quotes.toscrape.com&lt;/a&gt;). Playwright runs the scripts and gives us the real content:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/01-vscode-1-basics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/01-vscode-1-basics/</guid><description>&lt;h1 id="vs-code--part-1-basics"&gt;VS Code — Part 1: Basics&lt;a class="anchor" href="#vs-code--part-1-basics"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;VS Code is not just a text editor. It is a full development workspace: files, terminal, Git, debugger, extensions, and settings all integrated in one window. Most beginners treat it like Notepad. That is fine early on, but it limits you.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s what you will use VS Code for in real work:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Open project → write code → run in terminal → read errors
→ debug step by step → format code → commit to Git → repeat&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Each section below maps a VS Code feature to a real need in that cycle.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/01-vscode-2-advanced/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/01-vscode-2-advanced/</guid><description>&lt;h1 id="vs-code--part-2-python-debugging-git-and-workflow"&gt;VS Code — Part 2: Python, Debugging, Git, and Workflow&lt;a class="anchor" href="#vs-code--part-2-python-debugging-git-and-workflow"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This continues from Part 1. Make sure you have the &lt;code&gt;vscode-lab/&lt;/code&gt; folder open in VS Code with &lt;code&gt;app.py&lt;/code&gt;, &lt;code&gt;students.csv&lt;/code&gt;, &lt;code&gt;config.json&lt;/code&gt;, and &lt;code&gt;README.md&lt;/code&gt; created.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="python-setup-interpreter-first-packages-second"&gt;Python Setup: interpreter first, packages second&lt;a class="anchor" href="#python-setup-interpreter-first-packages-second"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Write this in &lt;code&gt;app.py&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; csv
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; json
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; open(&lt;span style="color:#e6db74"&gt;&amp;#34;config.json&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;r&amp;#34;&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; f:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; config &lt;span style="color:#f92672"&gt;=&lt;/span&gt; json&lt;span style="color:#f92672"&gt;.&lt;/span&gt;load(f)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;&amp;#34;App:&amp;#34;&lt;/span&gt;, config[&lt;span style="color:#e6db74"&gt;&amp;#34;appName&amp;#34;&lt;/span&gt;])
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;&amp;#34;Currency:&amp;#34;&lt;/span&gt;, config[&lt;span style="color:#e6db74"&gt;&amp;#34;currency&amp;#34;&lt;/span&gt;])
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;with&lt;/span&gt; open(&lt;span style="color:#e6db74"&gt;&amp;#34;students.csv&amp;#34;&lt;/span&gt;, &lt;span style="color:#e6db74"&gt;&amp;#34;r&amp;#34;&lt;/span&gt;) &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; f:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; reader &lt;span style="color:#f92672"&gt;=&lt;/span&gt; csv&lt;span style="color:#f92672"&gt;.&lt;/span&gt;DictReader(f)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; row &lt;span style="color:#f92672"&gt;in&lt;/span&gt; reader:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(row[&lt;span style="color:#e6db74"&gt;&amp;#34;name&amp;#34;&lt;/span&gt;], row[&lt;span style="color:#e6db74"&gt;&amp;#34;marks&amp;#34;&lt;/span&gt;], row[&lt;span style="color:#e6db74"&gt;&amp;#34;city&amp;#34;&lt;/span&gt;])&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Run it:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/02-uv-1-basics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/02-uv-1-basics/</guid><description>&lt;h1 id="uv--part-1-python-environment-management"&gt;uv — Part 1: Python Environment Management&lt;a class="anchor" href="#uv--part-1-python-environment-management"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;code&gt;uv&lt;/code&gt; is a fast, modern replacement for &lt;code&gt;pip&lt;/code&gt;, &lt;code&gt;venv&lt;/code&gt;, &lt;code&gt;pyenv&lt;/code&gt;, and &lt;code&gt;pipx&lt;/code&gt; — all rolled into one tool built by Astral. Instead of juggling multiple tools, you use &lt;code&gt;uv&lt;/code&gt; for almost everything Python-environment-related.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the comparison at a glance:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;What you need&lt;/th&gt;
					&lt;th&gt;Old way&lt;/th&gt;
					&lt;th&gt;With &lt;code&gt;uv&lt;/code&gt;&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Install Python&lt;/td&gt;
					&lt;td&gt;system installer / pyenv&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv python install 3.12&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Create virtual env&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;python -m venv .venv&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv venv&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Install packages&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pip install requests&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv pip install requests&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Project dependencies&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;requirements.txt&lt;/code&gt; manually&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pyproject.toml&lt;/code&gt; + &lt;code&gt;uv.lock&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Run one-off tools&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pipx run ruff&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uvx ruff&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Run a script&lt;/td&gt;
					&lt;td&gt;activate venv → python file.py&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;uv run file.py&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="install-uv"&gt;Install uv&lt;a class="anchor" href="#install-uv"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Linux/macOS&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;curl -LsSf https://astral.sh/uv/install.sh | sh
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Verify&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv --version&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;uv&lt;/code&gt; is a standalone binary — it doesn&amp;rsquo;t depend on Python being installed. You can install it on a fresh machine and let it manage Python for you.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/02-uv-2-advanced/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/02-uv-2-advanced/</guid><description>&lt;h1 id="uv--part-2-tools-advanced-workflows-and-best-practices"&gt;uv — Part 2: Tools, Advanced Workflows, and Best Practices&lt;a class="anchor" href="#uv--part-2-tools-advanced-workflows-and-best-practices"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;This continues from Part 1. You should already know &lt;code&gt;uv venv&lt;/code&gt;, &lt;code&gt;uv pip install&lt;/code&gt;, &lt;code&gt;uv init&lt;/code&gt;, and &lt;code&gt;uv add&lt;/code&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="requirementstxt-vs-pyprojecttoml--when-to-use-which"&gt;&lt;code&gt;requirements.txt&lt;/code&gt; vs &lt;code&gt;pyproject.toml&lt;/code&gt; — when to use which&lt;a class="anchor" href="#requirementstxt-vs-pyprojecttoml--when-to-use-which"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Situation&lt;/th&gt;
					&lt;th&gt;Use&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Working on an old project or cloning from GitHub&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;requirements.txt&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Your deployment platform expects &lt;code&gt;requirements.txt&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;requirements.txt&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Starting a new project&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pyproject.toml&lt;/code&gt; + &lt;code&gt;uv.lock&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Want reproducible, locked environments&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pyproject.toml&lt;/code&gt; + &lt;code&gt;uv.lock&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;May eventually publish as a package&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pyproject.toml&lt;/code&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For learning: understand &lt;code&gt;requirements.txt&lt;/code&gt; to work with existing projects. Use &lt;code&gt;pyproject.toml&lt;/code&gt; for everything you build yourself.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/03-bash-scripting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/03-bash-scripting/</guid><description>&lt;h1 id="bash-scripting--practical-notes"&gt;Bash Scripting — Practical Notes&lt;a class="anchor" href="#bash-scripting--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Bash is useful because every Linux server, Docker container, CI runner, and cloud VM gives you a shell. Python is better for large programs, but Bash excels at gluing commands together: download → filter → transform → run script → save logs → schedule daily job. The core idea: &lt;strong&gt;small commands + pipes + files + automation&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="1-how-the-terminal-reads-your-command"&gt;1. How the terminal reads your command&lt;a class="anchor" href="#1-how-the-terminal-reads-your-command"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When you type:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/04-git-github/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/04-git-github/</guid><description>&lt;h1 id="practical-git--github-notes"&gt;Practical Git + GitHub Notes&lt;a class="anchor" href="#practical-git--github-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Git is used by developers to &lt;strong&gt;track code changes, work safely in branches, collaborate through GitHub, review changes through PRs, and recover from mistakes&lt;/strong&gt;. Think of Git as your local project history machine, and GitHub as the hosted collaboration platform.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[Working directory&amp;lt;br/&amp;gt;files you edit] --&amp;gt;|git add| B[Staging area&amp;lt;br/&amp;gt;selected changes]
 B --&amp;gt;|git commit| C[Local Git repo&amp;lt;br/&amp;gt;history on laptop]
 C --&amp;gt;|git push| D[GitHub remote&amp;lt;br/&amp;gt;shared repo]
 D --&amp;gt;|git fetch / pull| C&lt;/pre&gt;&lt;p&gt;Git stores snapshots of your project, not just file versions. Branches are lightweight pointers to commits, which is why branching is fast and cheap.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/05-sqlite/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/05-sqlite/</guid><description>&lt;h1 id="sqlite--practical-notes"&gt;SQLite — Practical Notes&lt;a class="anchor" href="#sqlite--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;SQLite = &lt;strong&gt;one database inside one file&lt;/strong&gt;. No server to install, no process to manage — just a &lt;code&gt;.db&lt;/code&gt; file you can copy, query, and share.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;CSV / JSON SQLite .db file PostgreSQL
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;simple files → real queryable DB → big server DB
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;weak search local + portable multi-user production&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Use SQLite when you want:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;scraping data
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;course projects
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;local APIs
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;small dashboards
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;RAG metadata
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;logs / cache
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;searchable notes
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;quick demos&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Avoid SQLite when many machines/users must write to the same DB at the same time. Then use PostgreSQL.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/06-http-clients/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/06-http-clients/</guid><description>&lt;h1 id="http-clients--practical-notes"&gt;HTTP Clients — Practical Notes&lt;a class="anchor" href="#http-clients--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;h2 id="1-big-picture"&gt;1. Big Picture&lt;a class="anchor" href="#1-big-picture"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[Client: curl / Python / Browser] --&amp;gt; B[HTTP Request]
 B --&amp;gt; C[Server / API]
 C --&amp;gt; D[HTTP Response]
 D --&amp;gt; A

 B --&amp;gt; B1[Method: GET/POST/PATCH]
 B --&amp;gt; B2[URL]
 B --&amp;gt; B3[Headers]
 B --&amp;gt; B4[Body JSON/Form/File]

 D --&amp;gt; D1[Status Code]
 D --&amp;gt; D2[Headers]
 D --&amp;gt; D3[Body JSON/Text/File]&lt;/pre&gt;&lt;p&gt;HTTP client means: &lt;strong&gt;a tool or library used to talk to APIs&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Common clients:&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Client&lt;/th&gt;
					&lt;th&gt;Best use&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;curl&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Test/debug APIs from terminal&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Python &lt;code&gt;requests&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Simple Python API work&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Python &lt;code&gt;httpx&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Modern sync + async API work&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;httpbin&lt;/td&gt;
					&lt;td&gt;Safe API testing playground&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Postman / REST Client&lt;/td&gt;
					&lt;td&gt;GUI or project-based API testing&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="2-http-request-structure"&gt;2. HTTP Request Structure&lt;a class="anchor" href="#2-http-request-structure"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-txt" data-lang="txt"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;METHOD URL
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Headers
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Body&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Example:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/07-requestly/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/07-requestly/</guid><description>&lt;h1 id="requestly--burp-suite--practical-network-debugging-notes"&gt;Requestly + Burp Suite — practical network debugging notes&lt;a class="anchor" href="#requestly--burp-suite--practical-network-debugging-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Use these tools only on &lt;strong&gt;your own apps, local projects, staging systems, or targets where you have permission&lt;/strong&gt;. They are powerful because they can inspect and modify HTTP/HTTPS traffic.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[Browser] --&amp;gt; B{Debug need?}
 B --&amp;gt; C[Requestly]
 B --&amp;gt; D[Burp Suite]

 C --&amp;gt; C1[Quick browser rules]
 C --&amp;gt; C2[Redirect API]
 C --&amp;gt; C3[Modify headers]
 C --&amp;gt; C4[Mock response]

 D --&amp;gt; D1[Full proxy]
 D --&amp;gt; D2[Intercept request]
 D --&amp;gt; D3[Inspect HTTPS]
 D --&amp;gt; D4[Replay in Repeater]&lt;/pre&gt;&lt;h2 id="what-these-tools-are-used-for"&gt;What these tools are used for&lt;a class="anchor" href="#what-these-tools-are-used-for"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Requestly&lt;/strong&gt; is best for quick frontend/QA debugging directly in the browser: redirect URLs, modify headers/query params, mock API responses, and inject scripts/styles. Its browser extension works without setting up a local proxy, and Requestly describes it as useful for frontend debugging and simulating API responses.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/08-data-formats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/08-data-formats/</guid><description>&lt;h1 id="data-formats--practical-notes"&gt;Data Formats — Practical Notes&lt;a class="anchor" href="#data-formats--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Data formats are how programs &lt;strong&gt;store, send, read, and explain data&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In real developer work:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;API response -&amp;gt; JSON
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Python project -&amp;gt; TOML / pyproject.toml
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;GitHub Actions -&amp;gt; YAML
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;README / notes -&amp;gt; Markdown
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Image/PDF in JSON -&amp;gt; Base64
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Fast internal data -&amp;gt; MessagePack / Protobuf
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;LLM prompt data -&amp;gt; JSON / Markdown / TOON
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;All text -&amp;gt; UTF-8&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre class="mermaid"&gt;flowchart LR
 A[Data in program] --&amp;gt; B{Where is it going?}
 B --&amp;gt; C[API / web app: JSON]
 B --&amp;gt; D[Config: YAML or TOML]
 B --&amp;gt; E[Docs: Markdown]
 B --&amp;gt; F[Binary inside text: Base64]
 B --&amp;gt; G[High speed internal: MessagePack]
 B --&amp;gt; H[LLM context: JSON / TOON / Markdown]&lt;/pre&gt;&lt;h2 id="start-with-a-small-lab"&gt;Start with a small lab&lt;a class="anchor" href="#start-with-a-small-lab"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;mkdir data-formats-lab
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cd data-formats-lab
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Use uv if you have it&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv init --bare
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv add pyyaml msgpack tomli-w
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Optional but very useful for JSON&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sudo apt-get update
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sudo apt-get install -y jq
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Check tools&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python --version
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;jq --version&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="utf-8-the-rule-below-every-format"&gt;UTF-8: the rule below every format&lt;a class="anchor" href="#utf-8-the-rule-below-every-format"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Always assume text is &lt;strong&gt;Unicode&lt;/strong&gt;, and store/send it as &lt;strong&gt;UTF-8&lt;/strong&gt;. Unicode gives each character a number; UTF-8 turns that character into bytes. Normal ASCII characters use 1 byte, many non-Latin characters use 2–3 bytes, and emoji can use 4 bytes.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/09-github-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/09-github-pages/</guid><description>&lt;h1 id="github-pages--practical-notes"&gt;GitHub Pages — Practical Notes&lt;a class="anchor" href="#github-pages--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;GitHub Pages hosts &lt;strong&gt;static websites&lt;/strong&gt; directly from a GitHub repository. Static means the final output is only:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;HTML + CSS + JavaScript + images + fonts + JSON files&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;No backend server, no database server, no secret server-side environment variables. It is good for portfolios, documentation, project demos, blogs, landing pages, static dashboards, and built frontend apps. GitHub Pages supports free HTTPS and custom domains; project sites usually live at &lt;code&gt;https://&amp;lt;username&amp;gt;.github.io/&amp;lt;repo&amp;gt;/&lt;/code&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-1/10-latex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-1/10-latex/</guid><description>&lt;h1 id="latex--practical-notes"&gt;LaTeX — Practical Notes&lt;a class="anchor" href="#latex--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;LaTeX is used when Markdown is not enough: &lt;strong&gt;math equations, project reports, research-style PDFs, citations, figures, tables, and academic documents&lt;/strong&gt;. Pandoc can convert Markdown to LaTeX/PDF without writing raw &lt;code&gt;.tex&lt;/code&gt; files.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[Markdown .md] --&amp;gt;|pandoc| B[PDF via LaTeX]
 C[LaTeX .tex file] --&amp;gt;|pdflatex / latexmk| D[PDF]
 D --&amp;gt; E[polished report]&lt;/pre&gt;&lt;h2 id="use-latex-in-two-ways"&gt;Use LaTeX in two ways&lt;a class="anchor" href="#use-latex-in-two-ways"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;For beginners, start with &lt;strong&gt;Overleaf&lt;/strong&gt; (overleaf.com) — no install, browser-based editor, live PDF preview. For local or automated work, install LaTeX and Pandoc so reports can be built from the terminal or GitHub Actions.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/01-fastapi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/01-fastapi/</guid><description>&lt;h1 id="fastapi--practical-notes"&gt;FastAPI — Practical Notes&lt;a class="anchor" href="#fastapi--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;FastAPI is used when you want to turn Python code into an HTTP API. In real developer/data-science work, this means:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Python function → API endpoint → browser/app/script can call it&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Examples:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;train model locally → expose /predict
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;clean CSV → expose /clean
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summarize text → expose /summarize
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;long job → start it now, check result later
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cache expensive result → Redis&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre class="mermaid"&gt;flowchart LR
 A[Client: browser / curl / Python] --&amp;gt; B[FastAPI app]
 B --&amp;gt; C[Python function]
 C --&amp;gt; D[JSON response]
 B --&amp;gt; E[Redis cache / job status]
 B --&amp;gt; F[Background task]&lt;/pre&gt;&lt;p&gt;Start a project first.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/02-cors-middleware/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/02-cors-middleware/</guid><description>&lt;h2 id="browser-security--cors--middleware-in-fastapi"&gt;Browser security → CORS → Middleware in FastAPI&lt;a class="anchor" href="#browser-security--cors--middleware-in-fastapi"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When a frontend in the browser calls an API, the browser checks &lt;strong&gt;origin&lt;/strong&gt;.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Origin = scheme + domain + port
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;http://localhost:3000
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;https://localhost:3000
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;http://localhost:8000
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;https://api.example.com&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;These are different origins:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;http://localhost:3000 frontend
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;http://localhost:8000 backend API&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;So this frontend call is &lt;strong&gt;cross-origin&lt;/strong&gt;:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;fetch&lt;/span&gt;(&lt;span style="color:#e6db74"&gt;&amp;#34;http://localhost:8000/predict&amp;#34;&lt;/span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The browser has a security rule called &lt;strong&gt;Same-Origin Policy&lt;/strong&gt;. It stops JavaScript on one origin from freely reading sensitive data from another origin. CORS is the controlled way for a server to say: “This frontend origin is allowed to read my API response.” MDN defines CORS as an HTTP-header mechanism where the server indicates which other origins may access resources.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/03-google-oauth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/03-google-oauth/</guid><description>&lt;h2 id="google-oauth-with-fastapi-sessions-bearer-tokens-browser-storage"&gt;Google OAuth with FastAPI: sessions, bearer tokens, browser storage&lt;a class="anchor" href="#google-oauth-with-fastapi-sessions-bearer-tokens-browser-storage"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Before Google OAuth, understand two login styles.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[User logs in] --&amp;gt; B{How does server remember user?}
 B --&amp;gt; C[Session cookie]
 B --&amp;gt; D[Bearer token]&lt;/pre&gt;&lt;p&gt;A &lt;strong&gt;session&lt;/strong&gt; means the browser stores a small cookie. On every request, the browser automatically sends it back.&lt;/p&gt;
&lt;pre class="mermaid"&gt;sequenceDiagram
 participant B as Browser
 participant API as FastAPI

 B-&amp;gt;&amp;gt;API: POST /login
 API--&amp;gt;&amp;gt;B: Set-Cookie: session=abc...
 B-&amp;gt;&amp;gt;API: GET /me with Cookie
 API--&amp;gt;&amp;gt;B: user data&lt;/pre&gt;&lt;p&gt;In FastAPI/Starlette, &lt;code&gt;SessionMiddleware&lt;/code&gt; gives signed cookie-based sessions, and the session cookie is &lt;code&gt;HttpOnly&lt;/code&gt;, so JavaScript cannot read it directly.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/05-config-management/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/05-config-management/</guid><description>&lt;h1 id="config-management--practical-notes"&gt;Config Management — Practical Notes&lt;a class="anchor" href="#config-management--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Config management means &lt;strong&gt;same code, different values&lt;/strong&gt;. Your app code should not change between laptop, GitHub Actions, Hugging Face Spaces, staging, or production. Only environment variables change. This is the 12-factor idea: store config in the environment, not hardcoded in code.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 Code[&amp;#34;main.py / app.py&amp;lt;br/&amp;gt;same everywhere&amp;#34;]
 Local[&amp;#34;Local .env&amp;lt;br/&amp;gt;dev values&amp;#34;]
 Bash[&amp;#34;~/.bashrc&amp;lt;br/&amp;gt;machine-wide values&amp;#34;]
 GitHub[&amp;#34;GitHub Actions Secrets&amp;lt;br/&amp;gt;CI/CD values&amp;#34;]
 HF[&amp;#34;Hugging Face Space Secrets&amp;lt;br/&amp;gt;deployment values&amp;#34;]

 Local --&amp;gt; Code
 Bash --&amp;gt; Code
 GitHub --&amp;gt; Code
 HF --&amp;gt; Code&lt;/pre&gt;&lt;p&gt;Real developers use config management for database URLs, API keys, model names, debug mode, allowed CORS origins, logging level, ports, feature flags, and deployment secrets. The golden rule: &lt;strong&gt;commit code and &lt;code&gt;.env.example&lt;/code&gt;; never commit &lt;code&gt;.env&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/06-docker-compose/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/06-docker-compose/</guid><description>&lt;h1 id="podman-compose--practical-notes"&gt;Podman Compose — Practical Notes&lt;a class="anchor" href="#podman-compose--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Use containers when your app is not one process anymore. Real projects usually have &lt;strong&gt;API + database + frontend + admin UI + cache&lt;/strong&gt;. Podman runs each part in an isolated container; Compose starts/stops them together with one file.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 Dev[Developer laptop] --&amp;gt; Compose[compose.yml]
 Compose --&amp;gt; API[backend container]
 Compose --&amp;gt; FE[frontend container]
 Compose --&amp;gt; DB[postgres container]
 Compose --&amp;gt; UI[pgadmin container]
 DB --&amp;gt; Vol[(named volume)]
 API --&amp;gt; DB
 UI --&amp;gt; DB&lt;/pre&gt;&lt;p&gt;Podman is free and open source (Apache License). &lt;code&gt;podman-compose&lt;/code&gt; is a Compose implementation for Podman focused on rootless and daemonless use.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/07-deployment-platforms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/07-deployment-platforms/</guid><description>&lt;h1 id="deployment-platforms--practical-notes"&gt;Deployment Platforms — Practical Notes&lt;a class="anchor" href="#deployment-platforms--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Deployment means: your app stops living only on &lt;code&gt;localhost&lt;/code&gt; and starts running on a public URL.&lt;/p&gt;
&lt;p&gt;Real developer flow:&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 A[Code on laptop] --&amp;gt; B[Test locally]
 B --&amp;gt; C[Push to GitHub]
 C --&amp;gt; D[Platform builds app]
 D --&amp;gt; E[Platform runs app]
 E --&amp;gt; F[Public URL]
 F --&amp;gt; G[Logs, health checks, env vars]&lt;/pre&gt;&lt;p&gt;Use deployment platforms when you want to share APIs, demos, dashboards, ML apps, assignments, prototypes, or production services.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/08-logging-testing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/08-logging-testing/</guid><description>&lt;h1 id="python--fastapi-logging-and-testing--practical-notes"&gt;Python + FastAPI Logging and Testing — practical notes&lt;a class="anchor" href="#python--fastapi-logging-and-testing--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;In real developer work, &lt;strong&gt;logging&lt;/strong&gt; answers “what happened in the running app?” and &lt;strong&gt;testing&lt;/strong&gt; answers “does the code still work after changes?”&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 User[User / Browser / Client] --&amp;gt; API[FastAPI App]
 API --&amp;gt; Code[Business Logic]
 Code --&amp;gt; DB[(DB / File / Cache)]

 API --&amp;gt; Logs[Logs: what happened?]
 Code --&amp;gt; Tests[Tests: does it still work?]

 Logs --&amp;gt; Debug[Debug production issue]
 Tests --&amp;gt; SafeChange[Change code safely]&lt;/pre&gt;&lt;p&gt;Python’s standard &lt;code&gt;logging&lt;/code&gt; module is the normal built-in way to create loggers with &lt;code&gt;logging.getLogger(__name__)&lt;/code&gt;, and pytest is commonly used for small readable tests that scale to bigger apps. FastAPI’s &lt;code&gt;TestClient&lt;/code&gt; lets you test API endpoints using normal Python test functions and normal &lt;code&gt;assert&lt;/code&gt; statements.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/09-observability/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/09-observability/</guid><description>&lt;h1 id="observability--prometheus-with-pythonfastapi"&gt;Observability + Prometheus with Python/FastAPI&lt;a class="anchor" href="#observability--prometheus-with-pythonfastapi"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Observability means: &lt;strong&gt;when your app runs, you can understand what is happening from outside&lt;/strong&gt;. OpenTelemetry defines observability around telemetry data such as &lt;strong&gt;logs, metrics, and traces&lt;/strong&gt;; the application must be instrumented to emit this data. Prometheus is mainly for &lt;strong&gt;metrics&lt;/strong&gt;: it collects/scrapes metric values from HTTP endpoints, stores them as time series, and lets you query them with PromQL.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 User[User / Client] --&amp;gt; API[FastAPI App]

 API --&amp;gt; Logs[Logs&amp;lt;br/&amp;gt;What happened?]
 API --&amp;gt; Metrics[Metrics&amp;lt;br/&amp;gt;How much / how fast?]
 API --&amp;gt; Traces[Traces&amp;lt;br/&amp;gt;Where did time go?]

 Logs --&amp;gt; Debug[Debug errors]
 Metrics --&amp;gt; Alert[Alert on problems]
 Traces --&amp;gt; Slow[Find slow service/request]&lt;/pre&gt;&lt;p&gt;In simple words:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/10-cloudflare-tunnels/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/10-cloudflare-tunnels/</guid><description>&lt;h1 id="cloudflare-tunnel--other-tunnels--practical-developer-notes"&gt;Cloudflare Tunnel + other tunnels — practical developer notes&lt;a class="anchor" href="#cloudflare-tunnel--other-tunnels--practical-developer-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;A tunnel is used when your app is running on your laptop/server, but someone outside your network needs to access it. Example: webhook testing, demo to professor, mobile app testing, temporary API sharing, or exposing a home server safely.&lt;/p&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 Dev[Your laptop] --&amp;gt; App[FastAPI app&amp;lt;br/&amp;gt;localhost:8000]
 Friend[Friend / webhook / professor] -.cannot directly reach.-&amp;gt; App
 Friend --&amp;gt; PublicURL[Public tunnel URL]
 PublicURL --&amp;gt; TunnelProvider[Cloudflare / ngrok / Tailscale]
 TunnelProvider --&amp;gt; App&lt;/pre&gt;&lt;hr&gt;
&lt;h2 id="first-understand-localhost-127001-0000-lan-ip-public-ip"&gt;First understand localhost, 127.0.0.1, 0.0.0.0, LAN IP, public IP&lt;a class="anchor" href="#first-understand-localhost-127001-0000-lan-ip-public-ip"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When you run FastAPI normally:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/11-local-llms-1-basics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/11-local-llms-1-basics/</guid><description>&lt;h1 id="local-llm-tools--practical-notes"&gt;Local LLM Tools — Practical Notes&lt;a class="anchor" href="#local-llm-tools--practical-notes"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;hr&gt;
&lt;p&gt;&lt;a href="https://youtu.be/U8lGbSaCCYI?si=LvFNfoY0WmWnNwSJ"&gt;&lt;img src="https://img.youtube.com/vi/U8lGbSaCCYI/0.jpg" alt="Local LLM Tools Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="one-page-summary-of-all-tools"&gt;One-page summary of all tools&lt;a class="anchor" href="#one-page-summary-of-all-tools"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Tool&lt;/th&gt;
					&lt;th&gt;Best use&lt;/th&gt;
					&lt;th&gt;Interface&lt;/th&gt;
					&lt;th&gt;Best user&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;LM Studio&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;GUI-based local LLM usage, model search, chat, local API, document chat&lt;/td&gt;
					&lt;td&gt;GUI + CLI + API&lt;/td&gt;
					&lt;td&gt;Beginner, AI user, developer testing&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Ollama&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Fast local model running, CLI, local REST API, coding-tool integration&lt;/td&gt;
					&lt;td&gt;CLI + API + Docker&lt;/td&gt;
					&lt;td&gt;Developer prototyping&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;llama.cpp&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Deep control over GGUF models, quantization, CPU/GPU tuning, offline inference&lt;/td&gt;
					&lt;td&gt;CLI + server + C/C++ backend&lt;/td&gt;
					&lt;td&gt;Systems-minded developer&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;vLLM&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Production LLM serving, batching, high-throughput GPU inference&lt;/td&gt;
					&lt;td&gt;Python + server + Docker&lt;/td&gt;
					&lt;td&gt;Backend/ML engineer&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;MLX / mlx-lm&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Apple Silicon-native inference/fine-tuning&lt;/td&gt;
					&lt;td&gt;Python + CLI + Apple stack&lt;/td&gt;
					&lt;td&gt;Mac / Apple developer&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Simple choice:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/11-local-llms-2-lmstudio-ollama/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/11-local-llms-2-lmstudio-ollama/</guid><description>&lt;h1 id="6-lm-studio-best-beginner-gui-and-local-api-tool"&gt;6. LM Studio: best beginner GUI and local API tool&lt;a class="anchor" href="#6-lm-studio-best-beginner-gui-and-local-api-tool"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/OOCioZC4tk0?si=MICcoDpAbHlpeXcK"&gt;&lt;img src="https://img.youtube.com/vi/OOCioZC4tk0/0.jpg" alt="LM Studio Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="61-what-it-is"&gt;6.1 What it is&lt;a class="anchor" href="#61-what-it-is"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;LM Studio is a desktop app for downloading, managing, chatting with, and serving local models. It supports local model search/download, chat UI, MCP servers, OpenAI-like endpoints, prompt/config management, and document chat. It runs on macOS, Windows, and Linux.&lt;/p&gt;
&lt;p&gt;Use it when students need to &lt;strong&gt;see local AI working quickly&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="62-installation"&gt;6.2 Installation&lt;a class="anchor" href="#62-installation"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Download the app from the official LM Studio website.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/11-local-llms-3-llamacpp-vllm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/11-local-llms-3-llamacpp-vllm/</guid><description>&lt;h1 id="8-llamacpp-best-low-level-control-tool"&gt;8. llama.cpp: best low-level control tool&lt;a class="anchor" href="#8-llamacpp-best-low-level-control-tool"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/8F_5pdcD3HY?si=zQT8bgUXqXxnqAXr"&gt;&lt;img src="https://img.youtube.com/vi/8F_5pdcD3HY/0.jpg" alt="Llama cpp Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="81-what-it-is"&gt;8.1 What it is&lt;a class="anchor" href="#81-what-it-is"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;llama.cpp is a C/C++ inference engine for running LLMs locally, especially GGUF models. It supports many backends including Metal, CUDA, HIP/ROCm, Vulkan, SYCL, OpenCL, WebGPU, CPU backends, and more.&lt;/p&gt;
&lt;p&gt;Use llama.cpp when you want to understand what is really happening under the hood.&lt;/p&gt;
&lt;h2 id="82-installation"&gt;8.2 Installation&lt;a class="anchor" href="#82-installation"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;macOS/Linux with Homebrew:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;brew install llama.cpp&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Windows:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;winget install llama.cpp&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Build from source:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-2/11-local-llms-4-mlx-labs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-2/11-local-llms-4-mlx-labs/</guid><description>&lt;h1 id="10-mlx--mlx-lm-apple-silicon-native-path"&gt;10. MLX / mlx-lm: Apple Silicon-native path&lt;a class="anchor" href="#10-mlx--mlx-lm-apple-silicon-native-path"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/wykPErJ8M-8?si=dLvXpnoaLjCITTC4"&gt;&lt;img src="https://img.youtube.com/vi/wykPErJ8M-8/0.jpg" alt="MLX Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="101-what-it-is"&gt;10.1 What it is&lt;a class="anchor" href="#101-what-it-is"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;MLX is Apple’s machine-learning framework for Apple Silicon. It is designed for efficient and flexible machine learning on Apple Silicon, has Python/C++/Swift/C support, uses a NumPy-like API, supports CPU/GPU, and is designed around unified memory.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;mlx-lm&lt;/code&gt; is the LLM package for generating text and fine-tuning LLMs on Apple Silicon. It supports Hugging Face Hub integration, quantization, upload, low-rank/full fine-tuning, and distributed inference/fine-tuning.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/01-prompt-engineering-1-foundations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/01-prompt-engineering-1-foundations/</guid><description>&lt;h1 id="prompt-engineering--chapter-1-foundations"&gt;Prompt Engineering — Chapter 1: Foundations&lt;a class="anchor" href="#prompt-engineering--chapter-1-foundations"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=2BpCk4d2Cc0"&gt;&lt;img src="https://img.youtube.com/vi/2BpCk4d2Cc0/0.jpg" alt="Prompt Engineering Video 1" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=_ZvnD73m40o"&gt;&lt;img src="https://img.youtube.com/vi/_ZvnD73m40o/0.jpg" alt="Prompt Engineering Video 2" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Prompt engineering is not magic wording. It is the habit of giving an AI model a clear task, useful context, strong boundaries, and a verifiable output format.&lt;/p&gt;
&lt;p&gt;A good beginner rule:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Bad prompt = vague wish
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Good prompt = task + context + constraints + output format + examples when needed&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre class="mermaid"&gt;flowchart LR
 A[Your goal] --&amp;gt; B[Prompt]
 B --&amp;gt; C[LLM]
 C --&amp;gt; D[Output]
 D --&amp;gt; E{Good enough?}
 E --&amp;gt;|No| F[Add missing context / format / examples]
 F --&amp;gt; B
 E --&amp;gt;|Yes| G[Use or save prompt]&lt;/pre&gt;&lt;p&gt;Official docs from OpenAI and Anthropic agree on the core idea: be clear, specific, structured, and iterative. Claude additionally recommends XML tags heavily for complex prompts. OpenAI recommends putting important instructions first, separating context clearly, and using examples/output formats for reliability.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/01-prompt-engineering-2-reliable-reasoning-output-control/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/01-prompt-engineering-2-reliable-reasoning-output-control/</guid><description>&lt;h1 id="prompt-engineering--chapter-2-reliable-reasoning-and-output-control"&gt;Prompt Engineering — Chapter 2: Reliable Reasoning and Output Control&lt;a class="anchor" href="#prompt-engineering--chapter-2-reliable-reasoning-and-output-control"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Chapter 1 gave the basic prompt parts: task, context, constraints, examples, and output format. Chapter 2 is about making the model &lt;strong&gt;reason better&lt;/strong&gt;, &lt;strong&gt;avoid vague answers&lt;/strong&gt;, and &lt;strong&gt;return output you can actually use&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;A good mental model:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Beginner prompt = ask the model to answer
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Reliable prompt = ask the model to solve + check + format&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre class="mermaid"&gt;flowchart LR
 A[Task] --&amp;gt; B[Prompt with context]
 B --&amp;gt; C[Model reasons internally]
 C --&amp;gt; D[Draft answer]
 D --&amp;gt; E[Self-check against criteria]
 E --&amp;gt; F[Final structured answer]&lt;/pre&gt;&lt;p&gt;Do not chase hundreds of fancy prompting tricks. Most reliable prompting comes from five habits:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/01-prompt-engineering-3-prompted-applications-production-practice/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/01-prompt-engineering-3-prompted-applications-production-practice/</guid><description>&lt;h1 id="prompt-engineering--chapter-3-prompted-applications-and-production-practice"&gt;Prompt Engineering — Chapter 3: Prompted Applications and Production Practice&lt;a class="anchor" href="#prompt-engineering--chapter-3-prompted-applications-and-production-practice"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Chapter 1 taught prompt basics. Chapter 2 taught reliable reasoning and output control. Chapter 3 is where prompting becomes a &lt;strong&gt;real application&lt;/strong&gt;: the model uses documents, calls tools, follows safety boundaries, and is tested before changes go live.&lt;/p&gt;
&lt;p&gt;A normal prompt answers from model memory. A production prompt should answer from &lt;strong&gt;fresh context&lt;/strong&gt;, &lt;strong&gt;tools&lt;/strong&gt;, and &lt;strong&gt;rules your application controls&lt;/strong&gt;.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Chat prompt = user asks → model replies
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;AI application = user asks → retrieve/use tools/check rules → model replies&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre class="mermaid"&gt;flowchart LR
 A[User question] --&amp;gt; B[App layer]
 B --&amp;gt; C{Need knowledge?}
 C --&amp;gt;|Yes| D[Retrieve docs / search / DB]
 C --&amp;gt;|No| E[Use prompt only]
 B --&amp;gt; F{Need action?}
 F --&amp;gt;|Yes| G[Tool / API call]
 F --&amp;gt;|No| H[No tool]
 D --&amp;gt; I[Final prompt with evidence]
 G --&amp;gt; I
 E --&amp;gt; I
 H --&amp;gt; I
 I --&amp;gt; J[Model response]
 J --&amp;gt; K[Validate + log + return]&lt;/pre&gt;&lt;p&gt;The simple rule:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/ai-coding-assistants/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/ai-coding-assistants/</guid><description>&lt;h1 id="ai-coding-assistants"&gt;AI Coding Assistants&lt;a class="anchor" href="#ai-coding-assistants"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/gq5uGNSUVx4?si=aWwYbKD-rrObxDWO"&gt;&lt;img src="https://img.youtube.com/vi/gq5uGNSUVx4/0.jpg" alt="AI Coding Assistants Video 1" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/6C0FjHoN3qE?si=uZbMdP9FEMkc941u"&gt;&lt;img src="https://img.youtube.com/vi/6C0FjHoN3qE/0.jpg" alt="AI Coding Assistants Video 2" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;AI coding assistants are not just autocomplete. Used well, they can write, test, refactor, and explain entire codebases. Used poorly, they produce confident-sounding but wrong code. This page is about using them well.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="claude-code"&gt;Claude Code&lt;a class="anchor" href="#claude-code"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Claude Code is Anthropic&amp;rsquo;s agentic CLI coding tool. It reads your project files, writes code, runs commands, and iterates — all from the terminal.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/context-engineering/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/context-engineering/</guid><description>&lt;h1 id="context-engineering"&gt;Context Engineering&lt;a class="anchor" href="#context-engineering"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/jLuwLJBQkIs?si=MAApWvEI5gvVfVaH"&gt;&lt;img src="https://img.youtube.com/vi/jLuwLJBQkIs/0.jpg" alt="Context Engineering Video 1" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/4GiqzUHD5AA?si=QNpX4YmBqUbyKuUh"&gt;&lt;img src="https://img.youtube.com/vi/4GiqzUHD5AA/0.jpg" alt="Context Engineering Video 2" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Context engineering is the discipline of designing &lt;strong&gt;what goes in the context window&lt;/strong&gt; — the system prompt, background documents, instructions, tools, and conversation history — to make an LLM behave consistently and correctly across all inputs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prompt Engineering vs Context Engineering&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Prompt Engineering&lt;/strong&gt;: How to phrase one query to get a good answer&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context Engineering&lt;/strong&gt;: How to architect the entire context so the model behaves reliably &lt;em&gt;on every query&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="the-context-window-budget"&gt;The Context Window Budget&lt;a class="anchor" href="#the-context-window-budget"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Every LLM has a finite context window (measured in tokens). You decide how to spend it:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/langsmith-litellm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/langsmith-litellm/</guid><description>&lt;h1 id="langsmith--litellm"&gt;LangSmith &amp;amp; LiteLLM&lt;a class="anchor" href="#langsmith--litellm"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Two tools that give you &lt;strong&gt;visibility and control&lt;/strong&gt; over your LLM API usage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LangSmith&lt;/strong&gt; — trace every LLM call: see inputs, outputs, latency, cost per call&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LiteLLM&lt;/strong&gt; — one API to call 100+ LLM providers; built-in spend tracking and budget limits&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id="langsmith"&gt;LangSmith&lt;a class="anchor" href="#langsmith"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;LangSmith is an observability platform for LLM applications. Every call is logged as a trace — you see exactly what prompt went in, what came out, how long it took, and what it cost.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/llm-architecture-survey/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/llm-architecture-survey/</guid><description>&lt;h1 id="llm-architecture-survey"&gt;LLM Architecture Survey&lt;a class="anchor" href="#llm-architecture-survey"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;Understanding how LLMs work under the hood helps you use them better: why does temperature affect creativity? Why do longer prompts cost more? Why does CoT work? This topic gives you the mental models.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="the-transformer-2017--the-foundation-of-everything"&gt;The Transformer (2017 — The Foundation of Everything)&lt;a class="anchor" href="#the-transformer-2017--the-foundation-of-everything"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Every modern LLM is built on the Transformer architecture introduced in &amp;ldquo;Attention Is All You Need&amp;rdquo; (Vaswani et al., 2017).&lt;/p&gt;
&lt;h3 id="core-components"&gt;Core Components&lt;a class="anchor" href="#core-components"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Input text: &amp;#34;The cat sat on&amp;#34;
 ↓
[Tokenization] → [1192, 4690, 6654, 319] (token IDs)
 ↓
[Token Embedding] → 4 vectors of 768 dims each
 ↓
[Positional Encoding] → adds position information
 ↓
[Transformer Blocks] × N layers:
 │
 ├── [Self-Attention] — tokens attend to each other
 │ &amp;#34;sat&amp;#34; pays attention to &amp;#34;cat&amp;#34; (subject) and &amp;#34;on&amp;#34; (position)
 │
 ├── [Feed-Forward Network] — per-token transformation
 │ applies learned patterns to each position
 │
 └── [Layer Norm + Residual] — stability and gradient flow
 ↓
[Output Head] → probability over 50,000+ tokens
 ↓
Sample: &amp;#34;mat&amp;#34; (with probability 0.73)&lt;/code&gt;&lt;/pre&gt;&lt;h3 id="self-attention-the-core-mechanism"&gt;Self-Attention: The Core Mechanism&lt;a class="anchor" href="#self-attention-the-core-mechanism"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; numpy &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; np
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;scaled_dot_product_attention&lt;/span&gt;(Q, K, V, mask&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;):
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; Q: queries (seq_len × d_k)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; K: keys (seq_len × d_k)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; V: values (seq_len × d_v)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; d_k &lt;span style="color:#f92672"&gt;=&lt;/span&gt; Q&lt;span style="color:#f92672"&gt;.&lt;/span&gt;shape[&lt;span style="color:#f92672"&gt;-&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Step 1: compute attention scores&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; Q &lt;span style="color:#f92672"&gt;@&lt;/span&gt; K&lt;span style="color:#f92672"&gt;.&lt;/span&gt;T &lt;span style="color:#f92672"&gt;/&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sqrt(d_k) &lt;span style="color:#75715e"&gt;# scale to prevent vanishing gradients&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Step 2: optional masking (for decoder — can&amp;#39;t look at future tokens)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; mask &lt;span style="color:#f92672"&gt;is&lt;/span&gt; &lt;span style="color:#f92672"&gt;not&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; scores &lt;span style="color:#f92672"&gt;+&lt;/span&gt; mask &lt;span style="color:#f92672"&gt;*&lt;/span&gt; &lt;span style="color:#f92672"&gt;-&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1e9&lt;/span&gt; &lt;span style="color:#75715e"&gt;# -inf → 0 after softmax&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Step 3: softmax to get attention weights (sum to 1)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; weights &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;exp(scores) &lt;span style="color:#f92672"&gt;/&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sum(np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;exp(scores), axis&lt;span style="color:#f92672"&gt;=-&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;, keepdims&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Step 4: weighted sum of values&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; output &lt;span style="color:#f92672"&gt;=&lt;/span&gt; weights &lt;span style="color:#f92672"&gt;@&lt;/span&gt; V
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; output, weights
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Intuition: for each token, how much should I &amp;#34;attend&amp;#34; to other tokens?&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# High weight = &amp;#34;this token is very relevant to understanding me&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;Why attention is powerful:&lt;/strong&gt; &amp;ldquo;The animal didn&amp;rsquo;t cross the street because it was too tired.&amp;rdquo; What does &amp;ldquo;it&amp;rdquo; refer to? Attention lets the model look back at all previous tokens and learn that &amp;ldquo;it&amp;rdquo; → &amp;ldquo;animal&amp;rdquo; has high relevance.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/llm-cli-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/llm-cli-tools/</guid><description>&lt;h1 id="llm-cli-tools"&gt;LLM CLI Tools&lt;a class="anchor" href="#llm-cli-tools"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;CLI tools for LLMs bring AI into your existing shell workflows. Pipe text in, get structured answers out. No Python script needed for quick tasks.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="simon-willisons-llm-cli"&gt;Simon Willison&amp;rsquo;s &lt;code&gt;llm&lt;/code&gt; CLI&lt;a class="anchor" href="#simon-willisons-llm-cli"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;llm&lt;/code&gt; is the most popular CLI for querying language models from the terminal. It supports 50+ models via plugins and works seamlessly with Unix pipes.&lt;/p&gt;
&lt;h3 id="installation"&gt;Installation&lt;a class="anchor" href="#installation"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Install with UV (recommended)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uv tool install llm
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Or with pip&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;pip install llm
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Verify&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm --version&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="setting-up-api-keys"&gt;Setting Up API Keys&lt;a class="anchor" href="#setting-up-api-keys"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# OpenAI (default model)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm keys set openai
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# → Paste your sk-... key&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Anthropic&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm install llm-anthropic
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm keys set anthropic
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# → Paste your sk-ant-... key&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Google Gemini&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm install llm-gemini
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm keys set gemini
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# → Paste your AIza... key&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Ollama (local — no key needed)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm install llm-ollama
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Ollama must be running: ollama serve&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="basic-usage"&gt;Basic Usage&lt;a class="anchor" href="#basic-usage"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Simple query (uses default model)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm &lt;span style="color:#e6db74"&gt;&amp;#34;What is the difference between TCP and UDP?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Specify a model&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -m claude-sonnet-4-6 &lt;span style="color:#e6db74"&gt;&amp;#34;Explain Docker in one sentence.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -m gpt-4o-mini &lt;span style="color:#e6db74"&gt;&amp;#34;Summarize this: &lt;/span&gt;&lt;span style="color:#66d9ef"&gt;$(&lt;/span&gt;cat readme.md&lt;span style="color:#66d9ef"&gt;)&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -m gemini-2.0-flash &lt;span style="color:#e6db74"&gt;&amp;#34;What is 847 × 293?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Ollama (local, free)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -m ollama/llama3.2 &lt;span style="color:#e6db74"&gt;&amp;#34;Write a haiku about Python.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# List all available models&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm models list&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="the-power-unix-pipes"&gt;The Power: Unix Pipes&lt;a class="anchor" href="#the-power-unix-pipes"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Summarize a file&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat long_document.txt | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Summarize this in 5 bullet points&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Explain code&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat main.py | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Explain what this code does and find any bugs&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Fix a Python error&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;python script.py 2&amp;gt;&amp;amp;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt; | llm &lt;span style="color:#e6db74"&gt;&amp;#34;This is a Python error. What is the fix?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Translate a file&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat README.md | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Translate this to Hindi&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Generate commit messages&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;git diff --staged | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Write a concise git commit message for these changes&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Review PR diffs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;gh pr diff &lt;span style="color:#ae81ff"&gt;42&lt;/span&gt; | llm -m claude-sonnet-4-6 &lt;span style="color:#e6db74"&gt;&amp;#34;Review this PR. Focus on bugs and security issues.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Summarize logs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;tail -n &lt;span style="color:#ae81ff"&gt;100&lt;/span&gt; /var/log/app.log | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Summarize these logs. Flag any errors or anomalies.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Convert formats&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat data.csv | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Convert this CSV to a Markdown table&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Extract structured data&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat invoice.txt | llm &lt;span style="color:#e6db74"&gt;&amp;#34;Extract: vendor name, amount, date. Output as JSON.&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="system-prompts-and-templates"&gt;System Prompts and Templates&lt;a class="anchor" href="#system-prompts-and-templates"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Use a system prompt&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -s &lt;span style="color:#e6db74"&gt;&amp;#34;You are a senior Python code reviewer. Be direct and critical.&amp;#34;&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;$(&lt;/span&gt;cat main.py&lt;span style="color:#66d9ef"&gt;)&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Save a template&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm templates edit code-reviewer&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Template format (YAML in your editor)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;system&lt;/span&gt;: |&lt;span style="color:#e6db74"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; You are an expert code reviewer. For each issue you find, provide:
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; 1. The line number
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; 2. The problem
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; 3. The fix&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;model&lt;/span&gt;: &lt;span style="color:#ae81ff"&gt;claude-sonnet-4-6&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Use the template&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat main.py | llm -t code-reviewer
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Create a &amp;#34;commit message&amp;#34; template&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm templates edit commit-msg
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# system: &amp;#34;Write a concise, imperative git commit message. Max 72 chars for subject line.&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="conversations"&gt;Conversations&lt;a class="anchor" href="#conversations"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Start a conversation (maintains history in session)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm chat -m claude-sonnet-4-6
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Or continue the last conversation&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -c &lt;span style="color:#e6db74"&gt;&amp;#34;What was the last thing I asked about?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Named conversation&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -s &lt;span style="color:#e6db74"&gt;&amp;#34;You are a Python tutor&amp;#34;&lt;/span&gt; --conversation my-python-session &lt;span style="color:#e6db74"&gt;&amp;#34;What is a generator?&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm -c --conversation my-python-session &lt;span style="color:#e6db74"&gt;&amp;#34;Give me an example.&amp;#34;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="logging-and-history"&gt;Logging and History&lt;a class="anchor" href="#logging-and-history"&gt;#&lt;/a&gt;&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# All queries are logged by default&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm logs &lt;span style="color:#75715e"&gt;# show recent queries&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm logs -n &lt;span style="color:#ae81ff"&gt;20&lt;/span&gt; &lt;span style="color:#75715e"&gt;# last 20&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm logs --json &lt;span style="color:#75715e"&gt;# as JSON&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;llm logs --json | jq &lt;span style="color:#e6db74"&gt;&amp;#39;.[] | {prompt, response}&amp;#39;&lt;/span&gt; &lt;span style="color:#75715e"&gt;# pipe to jq&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# The logs database is at:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# ~/.config/io.datasette.llm/logs.db&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Query it with datasette&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;datasette ~/.config/io.datasette.llm/logs.db&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;hr&gt;
&lt;h2 id="aichat"&gt;aichat&lt;a class="anchor" href="#aichat"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;aichat&lt;/code&gt; is a feature-rich terminal AI client with a built-in shell integration mode.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/multimodal-inputs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/multimodal-inputs/</guid><description>&lt;h1 id="multimodal-inputs"&gt;Multimodal Inputs&lt;a class="anchor" href="#multimodal-inputs"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/wdeZIOJjjdE?si=wFK2fsXu96TbYIiL"&gt;&lt;img src="https://img.youtube.com/vi/wdeZIOJjjdE/0.jpg" alt="Multimodal Inputs Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Modern LLMs aren&amp;rsquo;t just text-in, text-out. You can send images, PDFs, and audio directly in API requests and get back rich, structured analysis. This is the foundation of document processing pipelines, visual QA, and OCR-free data extraction.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="how-multimodal-apis-work"&gt;How Multimodal APIs Work&lt;a class="anchor" href="#how-multimodal-apis-work"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The API accepts a &lt;code&gt;content&lt;/code&gt; list instead of a plain string. Each item in the list is a &lt;strong&gt;content block&lt;/strong&gt; — either text, image, or document:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/prompt-caching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/prompt-caching/</guid><description>&lt;h1 id="prompt-caching"&gt;Prompt Caching&lt;a class="anchor" href="#prompt-caching"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/tECAkJAI_Vk?si=7zJK7B2nEtO5kD3S"&gt;&lt;img src="https://img.youtube.com/vi/tECAkJAI_Vk/0.jpg" alt="Prompt Caching Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Every time you call an LLM API, you pay to process the &lt;strong&gt;entire&lt;/strong&gt; prompt — even if the system prompt, the 50-page document, or the tool definitions haven&amp;rsquo;t changed since the last call. Prompt caching fixes this: process once, read from cache for a fraction of the cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real numbers (Anthropic Claude Sonnet 4.6)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Standard input: $3.00 / MTok&lt;/li&gt;
&lt;li&gt;Cache write: $3.75 / MTok (1.25×)&lt;/li&gt;
&lt;li&gt;Cache read: $0.30 / MTok (0.10×)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;A 10,000-token system prompt read 100 times:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/similarity-search/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/similarity-search/</guid><description>&lt;h1 id="similarity-search"&gt;Similarity Search&lt;a class="anchor" href="#similarity-search"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/nbJVJ1RPBEg?si=KTCsKw7zSdMcfBdt"&gt;&lt;img src="https://img.youtube.com/vi/nbJVJ1RPBEg/0.jpg" alt="Similarity Search Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You have a million embeddings. A user queries arrive. Brute-force cosine similarity over 1M × 1536-dim vectors takes seconds. &lt;strong&gt;Approximate Nearest Neighbor (ANN)&lt;/strong&gt; indices make it milliseconds — with only a small loss in recall.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="the-problem-brute-force-doesnt-scale"&gt;The Problem: Brute-Force Doesn&amp;rsquo;t Scale&lt;a class="anchor" href="#the-problem-brute-force-doesnt-scale"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; numpy &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; np
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; time
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;n_docs &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1_000_000&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;dim &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1536&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Simulate 1M document embeddings&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;corpus &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;random&lt;span style="color:#f92672"&gt;.&lt;/span&gt;randn(n_docs, dim)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;astype(np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;float32)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;corpus &lt;span style="color:#f92672"&gt;/=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;linalg&lt;span style="color:#f92672"&gt;.&lt;/span&gt;norm(corpus, axis&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;, keepdims&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# normalize&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;random&lt;span style="color:#f92672"&gt;.&lt;/span&gt;randn(dim)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;astype(np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;float32)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query &lt;span style="color:#f92672"&gt;/=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;linalg&lt;span style="color:#f92672"&gt;.&lt;/span&gt;norm(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Brute force&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;start &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;similarities &lt;span style="color:#f92672"&gt;=&lt;/span&gt; corpus &lt;span style="color:#f92672"&gt;@&lt;/span&gt; query
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;top5 &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;argsort(similarities)[::&lt;span style="color:#f92672"&gt;-&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;][:&lt;span style="color:#ae81ff"&gt;5&lt;/span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;elapsed &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time() &lt;span style="color:#f92672"&gt;-&lt;/span&gt; start
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Brute force: &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;elapsed&lt;span style="color:#f92672"&gt;*&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1000&lt;/span&gt;&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.0f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;ms for &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;n_docs&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;,&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt; docs&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# → ~1,200ms on CPU — too slow for a web API&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;ANN indices trade a small amount of accuracy for 100–1000x speed:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/structured-output/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/structured-output/</guid><description>&lt;h1 id="pydantic--structured-output"&gt;Pydantic &amp;amp; Structured Output&lt;a class="anchor" href="#pydantic--structured-output"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/fuMKrKlaku4?si=y1Tgb09oBW67CPVB"&gt;&lt;img src="https://img.youtube.com/vi/fuMKrKlaku4/0.jpg" alt="Structured Output Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;LLMs return free-form text. Your application needs typed, validated data. &lt;strong&gt;Instructor&lt;/strong&gt; is the bridge — it wraps any LLM client and guarantees structured, validated Pydantic output with automatic retry on failure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why this matters&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Without structured output, you write brittle regex parsers that break on every model update. With Instructor + Pydantic, you define a schema once and the LLM fills it.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-3/vector-embeddings/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-3/vector-embeddings/</guid><description>&lt;h1 id="vector-embeddings"&gt;Vector Embeddings&lt;a class="anchor" href="#vector-embeddings"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://youtu.be/9rYOLzoDNrM?si=0LQqYfjmf258U11a"&gt;&lt;img src="https://img.youtube.com/vi/9rYOLzoDNrM/0.jpg" alt="Vector Embeddings Video" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;An &lt;strong&gt;embedding&lt;/strong&gt; is a fixed-length list of numbers (a vector) that represents the meaning of a text. Texts with similar meanings have vectors that are close to each other in high-dimensional space. This is the mathematical foundation of semantic search, RAG, and recommendation systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intuition&lt;/strong&gt;
&amp;ldquo;King&amp;rdquo; − &amp;ldquo;Man&amp;rdquo; + &amp;ldquo;Woman&amp;rdquo; ≈ &amp;ldquo;Queen&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t magic — embeddings place semantically related words in similar geometric regions of the vector space.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/chunking-strategies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/chunking-strategies/</guid><description>&lt;h1 id="chunking-strategies"&gt;Chunking Strategies&lt;a class="anchor" href="#chunking-strategies"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The most underrated step in RAG.&lt;/strong&gt; Better chunking → better retrieval → better answers. No amount of fancy reranking rescues bad chunks.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="why-chunking-matters"&gt;Why Chunking Matters&lt;a class="anchor" href="#why-chunking-matters"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Your embedding model has a token limit (typically 512–8192 tokens). You can&amp;rsquo;t embed an entire PDF at once. But chunks that are too small lose context; chunks too large dilute the signal.&lt;/p&gt;
&lt;p&gt;The goal: &lt;strong&gt;chunks that are semantically coherent and self-contained enough to answer a question on their own.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/contextual-retrieval/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/contextual-retrieval/</guid><description>&lt;h1 id="contextual-retrieval"&gt;Contextual Retrieval&lt;a class="anchor" href="#contextual-retrieval"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;The problem: chunks lose context when extracted. The fix: prepend a custom context sentence to each chunk before embedding.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-core-idea"&gt;The Core Idea&lt;a class="anchor" href="#the-core-idea"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Anthropic published this technique in September 2024. The insight is deceptively simple:&lt;/p&gt;
&lt;p&gt;Before embedding a chunk, prepend a short &lt;strong&gt;context sentence&lt;/strong&gt; that describes where the chunk comes from and what it&amp;rsquo;s about — generated by a fast LLM using the full document as input.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/graphrag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/graphrag/</guid><description>&lt;h1 id="graphrag"&gt;GraphRAG&lt;a class="anchor" href="#graphrag"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Vector search finds similar passages. GraphRAG finds patterns, relationships, and themes across the entire corpus.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;a href="https://youtu.be/r09tJfON6kE" title="GraphRAG Explained — Microsoft Research"&gt;&lt;img src="https://img.youtube.com/vi/r09tJfON6kE/0.jpg" alt="GraphRAG Explained — Microsoft Research" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="why-vector-search-fails-at-global-questions"&gt;Why Vector Search Fails at Global Questions&lt;a class="anchor" href="#why-vector-search-fails-at-global-questions"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Standard RAG answers &amp;ldquo;local&amp;rdquo; questions well — questions answered by one or a few passages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It struggles with:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;What are the main themes across all documents?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;How do these companies relate to each other?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;What is the overall sentiment trend in these reports?&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These require &lt;strong&gt;synthesizing&lt;/strong&gt; information across the entire corpus — not just finding a relevant chunk.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/hybrid-search/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/hybrid-search/</guid><description>&lt;h1 id="hybrid-search"&gt;Hybrid Search&lt;a class="anchor" href="#hybrid-search"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Dense search finds synonyms. Sparse search finds exact terms. Hybrid search wins.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-problem-with-pure-vector-search"&gt;The Problem with Pure Vector Search&lt;a class="anchor" href="#the-problem-with-pure-vector-search"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Vector (dense) search is powerful but has blind spots:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Keyword mismatch&lt;/strong&gt; — &amp;ldquo;myocardial infarction&amp;rdquo; ≠ &amp;ldquo;heart attack&amp;rdquo; in dense space without enough training data&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rare terms&lt;/strong&gt; — product codes, model numbers, names get diluted&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exact match needs&lt;/strong&gt; — legal terms, codes, IDs should match exactly&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;BM25 (sparse) search has the opposite profile: great for exact terms, poor for semantic similarity.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/late-chunking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/late-chunking/</guid><description>&lt;h1 id="late-chunking"&gt;Late Chunking&lt;a class="anchor" href="#late-chunking"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Traditional chunking: chunk → embed. Late chunking: embed → chunk. That order reversal is everything.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-problem-with-early-chunking"&gt;The Problem with Early Chunking&lt;a class="anchor" href="#the-problem-with-early-chunking"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;When you chunk a document before embedding, each chunk loses context from the rest of the document.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt;&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Document: &amp;#34;The Eiffel Tower was built in 1889. It is located in Paris.
 This famous landmark attracts millions of visitors annually.&amp;#34;

Chunk 1: &amp;#34;The Eiffel Tower was built in 1889.&amp;#34;
Chunk 2: &amp;#34;It is located in Paris.&amp;#34; ← &amp;#34;It&amp;#34; has no referent!
Chunk 3: &amp;#34;This famous landmark attracts...&amp;#34; ← What landmark?&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Chunk 2&amp;rsquo;s embedding doesn&amp;rsquo;t know &amp;ldquo;It&amp;rdquo; refers to the Eiffel Tower. The embedding is incomplete.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/llm-grounding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/llm-grounding/</guid><description>&lt;h1 id="llm-grounding--citations"&gt;LLM Grounding &amp;amp; Citations&lt;a class="anchor" href="#llm-grounding--citations"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;An answer without a source is just an opinion. Ground your LLM to the documents it retrieved.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-grounding"&gt;What Is Grounding?&lt;a class="anchor" href="#what-is-grounding"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Grounding means forcing the LLM to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Only use retrieved context&lt;/strong&gt; — no hallucinated facts&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cite every claim&lt;/strong&gt; — map each sentence to a source chunk&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Admit uncertainty&lt;/strong&gt; — say &amp;ldquo;I don&amp;rsquo;t know&amp;rdquo; when context is insufficient&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="system-prompt-for-grounded-answers"&gt;System Prompt for Grounded Answers&lt;a class="anchor" href="#system-prompt-for-grounded-answers"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The simplest form of grounding is a strict system prompt:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/multimodal-embeddings/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/multimodal-embeddings/</guid><description>&lt;h1 id="multimodal-embeddings"&gt;Multimodal Embeddings&lt;a class="anchor" href="#multimodal-embeddings"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Text-only RAG misses 80% of real-world documents — PDFs with charts, tables, figures. Multimodal embeddings fix that.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-problem-with-text-only-rag"&gt;The Problem with Text-Only RAG&lt;a class="anchor" href="#the-problem-with-text-only-rag"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Standard RAG pipeline:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Parse PDF → extract text → chunk text → embed text&lt;/li&gt;
&lt;li&gt;Lost: tables, diagrams, headers in images, scanned pages&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Multimodal embedding&lt;/strong&gt; treats the document as an &lt;strong&gt;image&lt;/strong&gt; — no text extraction needed.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="clip--connecting-text-and-images"&gt;CLIP — Connecting Text and Images&lt;a class="anchor" href="#clip--connecting-text-and-images"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;CLIP (OpenAI) embeds images and text into the &lt;strong&gt;same vector space&lt;/strong&gt;. You can query with text and retrieve images, or vice versa.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/query-augmentation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/query-augmentation/</guid><description>&lt;h1 id="query-augmentation"&gt;Query Augmentation&lt;a class="anchor" href="#query-augmentation"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;The query the user types is rarely the best query for retrieval. Fix it before it hits your vector database.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="why-augment-queries"&gt;Why Augment Queries?&lt;a class="anchor" href="#why-augment-queries"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Users ask vague, short, or jargon-heavy questions. Your retrieval system works best with precise, information-rich queries.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;User Query&lt;/th&gt;
					&lt;th&gt;What They Mean&lt;/th&gt;
					&lt;th&gt;Better Retrieval Query&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;How does it work?&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;(no context)&lt;/td&gt;
					&lt;td&gt;Depends on the document!&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Python error in loop&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;IndexError in for loop&lt;/td&gt;
					&lt;td&gt;&amp;ldquo;Python IndexError list index out of range for loop&amp;rdquo;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&amp;ldquo;Fast RAG&amp;rdquo;&lt;/td&gt;
					&lt;td&gt;Low-latency RAG pipelines&lt;/td&gt;
					&lt;td&gt;&amp;ldquo;Techniques to reduce latency in retrieval-augmented generation&amp;rdquo;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="technique-1-hyde--hypothetical-document-embeddings"&gt;Technique 1: HyDE — Hypothetical Document Embeddings&lt;a class="anchor" href="#technique-1-hyde--hypothetical-document-embeddings"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Idea:&lt;/strong&gt; Instead of embedding the query, generate a &lt;em&gt;hypothetical&lt;/em&gt; answer document, then embed that.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/ragas-evaluation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/ragas-evaluation/</guid><description>&lt;h1 id="ragas-evaluation"&gt;RAGAS Evaluation&lt;a class="anchor" href="#ragas-evaluation"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;You can&amp;rsquo;t improve what you can&amp;rsquo;t measure.&amp;rdquo;&lt;/strong&gt; RAGAS gives you numbers, not vibes.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;a href="https://youtu.be/IlNglM9bKLw" title="RAGAS: Automated Evaluation of RAG Pipelines"&gt;&lt;img src="https://img.youtube.com/vi/IlNglM9bKLw/0.jpg" alt="RAGAS: Automated Evaluation of RAG Pipelines" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="why-evaluate"&gt;Why Evaluate?&lt;a class="anchor" href="#why-evaluate"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;RAG systems fail in subtle ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Retrieved the right documents but &lt;strong&gt;hallucinated&lt;/strong&gt; the answer&lt;/li&gt;
&lt;li&gt;Gave a correct-looking answer that &lt;strong&gt;wasn&amp;rsquo;t supported&lt;/strong&gt; by context&lt;/li&gt;
&lt;li&gt;Retrieved &lt;strong&gt;irrelevant&lt;/strong&gt; chunks despite having good documents in the corpus&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;RAGAS catches all three failure modes with four key metrics.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/reranking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/reranking/</guid><description>&lt;h1 id="reranking"&gt;Reranking&lt;a class="anchor" href="#reranking"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Retrieve many, rerank few, pass the best to the LLM.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-two-stage-retrieval-pattern"&gt;The Two-Stage Retrieval Pattern&lt;a class="anchor" href="#the-two-stage-retrieval-pattern"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;First-stage retrieval (vector search) is fast but approximate. Rerankers are slow but accurate. The trick: run the fast stage to get 20-50 candidates, then run the accurate stage on just those candidates.&lt;/p&gt;
&lt;pre class="mermaid"&gt;graph LR
 A[Query] --&amp;gt; B[Stage 1: ANN Search]
 B --&amp;gt; C[&amp;#34;Top-50 Candidates\n(fast, approximate)&amp;#34;]
 C --&amp;gt; D[Stage 2: Reranker]
 D --&amp;gt; E[&amp;#34;Top-5 Results\n(slow, accurate)&amp;#34;]
 E --&amp;gt; F[LLM]&lt;/pre&gt;&lt;p&gt;Cost: vector search on 1M docs + reranking on 50 docs &amp;raquo; vector search alone on 1M docs.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/semantic-caching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/semantic-caching/</guid><description>&lt;h1 id="semantic-caching"&gt;Semantic Caching&lt;a class="anchor" href="#semantic-caching"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;If two questions mean the same thing, why call the LLM twice?&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="the-problem"&gt;The Problem&lt;a class="anchor" href="#the-problem"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;In production RAG, many user questions are semantically identical:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;What is RAG?&amp;rdquo; / &amp;ldquo;Explain RAG&amp;rdquo; / &amp;ldquo;How does RAG work?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Who is the CEO?&amp;rdquo; / &amp;ldquo;Who leads the company?&amp;rdquo; / &amp;ldquo;Who runs this org?&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Standard caching (exact string match) misses all these. Semantic caching catches them.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="how-semantic-caching-works"&gt;How Semantic Caching Works&lt;a class="anchor" href="#how-semantic-caching-works"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;graph LR
 A[New Query] --&amp;gt; B[Embed Query]
 B --&amp;gt; C{Check Cache\nCosine Similarity}
 C --&amp;gt;|Hit: sim &amp;gt; threshold| D[Return Cached Answer]
 C --&amp;gt;|Miss: sim &amp;lt; threshold| E[Run RAG Pipeline]
 E --&amp;gt; F[Store in Cache\nquery embedding + answer]
 F --&amp;gt; G[Return Fresh Answer]&lt;/pre&gt;&lt;ol&gt;
&lt;li&gt;Embed the incoming query&lt;/li&gt;
&lt;li&gt;Search the cache for similar past queries&lt;/li&gt;
&lt;li&gt;If similarity &amp;gt; threshold (e.g., 0.92) → return cached answer&lt;/li&gt;
&lt;li&gt;If miss → run full RAG → cache the result&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2 id="build-a-semantic-cache-from-scratch"&gt;Build a Semantic Cache from Scratch&lt;a class="anchor" href="#build-a-semantic-cache-from-scratch"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; numpy &lt;span style="color:#66d9ef"&gt;as&lt;/span&gt; np
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; json
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; time
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; dataclasses &lt;span style="color:#f92672"&gt;import&lt;/span&gt; dataclass, asdict
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;from&lt;/span&gt; openai &lt;span style="color:#f92672"&gt;import&lt;/span&gt; OpenAI
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;client &lt;span style="color:#f92672"&gt;=&lt;/span&gt; OpenAI()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#a6e22e"&gt;@dataclass&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;class&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;CacheEntry&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; query: str
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; answer: str
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; embedding: list[float]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; timestamp: float
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; hit_count: int &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;class&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;SemanticCache&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; In-memory semantic cache with cosine similarity lookup.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; For production: replace the list with a vector database.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;__init__&lt;/span&gt;(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; similarity_threshold: float &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;0.92&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; embedding_model: str &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;text-embedding-3-small&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; max_size: int &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1000&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ttl_seconds: float &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;3600&lt;/span&gt;, &lt;span style="color:#75715e"&gt;# 1 hour TTL&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ):
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;threshold &lt;span style="color:#f92672"&gt;=&lt;/span&gt; similarity_threshold
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;model &lt;span style="color:#f92672"&gt;=&lt;/span&gt; embedding_model
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;max_size &lt;span style="color:#f92672"&gt;=&lt;/span&gt; max_size
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;ttl &lt;span style="color:#f92672"&gt;=&lt;/span&gt; ttl_seconds
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache: list[CacheEntry] &lt;span style="color:#f92672"&gt;=&lt;/span&gt; []
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;_embed&lt;/span&gt;(self, text: str) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; list[float]:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; response &lt;span style="color:#f92672"&gt;=&lt;/span&gt; client&lt;span style="color:#f92672"&gt;.&lt;/span&gt;embeddings&lt;span style="color:#f92672"&gt;.&lt;/span&gt;create(input&lt;span style="color:#f92672"&gt;=&lt;/span&gt;[text], model&lt;span style="color:#f92672"&gt;=&lt;/span&gt;self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;model)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; response&lt;span style="color:#f92672"&gt;.&lt;/span&gt;data[&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;]&lt;span style="color:#f92672"&gt;.&lt;/span&gt;embedding
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;_cosine_similarity&lt;/span&gt;(self, a: list[float], b: list[float]) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; float:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; a, b &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;array(a), np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;array(b)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; float(np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;dot(a, b) &lt;span style="color:#f92672"&gt;/&lt;/span&gt; (np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;linalg&lt;span style="color:#f92672"&gt;.&lt;/span&gt;norm(a) &lt;span style="color:#f92672"&gt;*&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;linalg&lt;span style="color:#f92672"&gt;.&lt;/span&gt;norm(b)))
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;_is_expired&lt;/span&gt;(self, entry: CacheEntry) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; bool:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; (time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time() &lt;span style="color:#f92672"&gt;-&lt;/span&gt; entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;timestamp) &lt;span style="color:#f92672"&gt;&amp;gt;&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;ttl
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;get&lt;/span&gt;(self, query: str) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; tuple[str &lt;span style="color:#f92672"&gt;|&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;, float]:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; Look up query in cache.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; Returns (answer, similarity) or (None, 0.0) on miss.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; query_embedding &lt;span style="color:#f92672"&gt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;_embed(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; best_similarity &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;0.0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; best_entry &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; now &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; entry &lt;span style="color:#f92672"&gt;in&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Skip expired entries&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; (now &lt;span style="color:#f92672"&gt;-&lt;/span&gt; entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;timestamp) &lt;span style="color:#f92672"&gt;&amp;gt;&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;ttl:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;continue&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; sim &lt;span style="color:#f92672"&gt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;_cosine_similarity(query_embedding, entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;embedding)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; sim &lt;span style="color:#f92672"&gt;&amp;gt;&lt;/span&gt; best_similarity:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; best_similarity &lt;span style="color:#f92672"&gt;=&lt;/span&gt; sim
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; best_entry &lt;span style="color:#f92672"&gt;=&lt;/span&gt; entry
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; best_entry &lt;span style="color:#f92672"&gt;and&lt;/span&gt; best_similarity &lt;span style="color:#f92672"&gt;&amp;gt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;threshold:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; best_entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;hit_count &lt;span style="color:#f92672"&gt;+=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;🎯 Cache HIT (sim=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;best_similarity&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.4f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;): &amp;#39;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;best_entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;query&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#39;&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; best_entry&lt;span style="color:#f92672"&gt;.&lt;/span&gt;answer, best_similarity
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;❌ Cache MISS (best sim=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;best_similarity&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.4f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;)&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;, best_similarity
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;set&lt;/span&gt;(self, query: str, answer: str) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;Store a query-answer pair in the cache.&amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; len(self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache) &lt;span style="color:#f92672"&gt;&amp;gt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;max_size:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Evict oldest entry (LRU-style)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;pop(&lt;span style="color:#ae81ff"&gt;0&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; embedding &lt;span style="color:#f92672"&gt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;_embed(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; entry &lt;span style="color:#f92672"&gt;=&lt;/span&gt; CacheEntry(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; query&lt;span style="color:#f92672"&gt;=&lt;/span&gt;query,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; answer&lt;span style="color:#f92672"&gt;=&lt;/span&gt;answer,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; embedding&lt;span style="color:#f92672"&gt;=&lt;/span&gt;embedding,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; timestamp&lt;span style="color:#f92672"&gt;=&lt;/span&gt;time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time(),
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; )
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;append(entry)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;✅ Cached: &amp;#39;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;query[:&lt;span style="color:#ae81ff"&gt;50&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;...&amp;#39;&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;stats&lt;/span&gt;(self) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; dict:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; now &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; active &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [e &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; e &lt;span style="color:#f92672"&gt;in&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; (now &lt;span style="color:#f92672"&gt;-&lt;/span&gt; e&lt;span style="color:#f92672"&gt;.&lt;/span&gt;timestamp) &lt;span style="color:#f92672"&gt;&amp;lt;=&lt;/span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;ttl]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; total_hits &lt;span style="color:#f92672"&gt;=&lt;/span&gt; sum(e&lt;span style="color:#f92672"&gt;.&lt;/span&gt;hit_count &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; e &lt;span style="color:#f92672"&gt;in&lt;/span&gt; active)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;total_entries&amp;#34;&lt;/span&gt;: len(self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache),
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;active_entries&amp;#34;&lt;/span&gt;: len(active),
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;total_cache_hits&amp;#34;&lt;/span&gt;: total_hits,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;threshold&amp;#34;&lt;/span&gt;: self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;threshold,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;clear&lt;/span&gt;(self) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;None&lt;/span&gt;:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; self&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;clear()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# --- Usage with RAG Pipeline ---&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;rag_with_cache&lt;/span&gt;(
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; query: str,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; cache: SemanticCache,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; rag_fn, &lt;span style="color:#75715e"&gt;# your actual RAG function&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; dict:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; Wrap any RAG pipeline with semantic caching.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; Returns answer + metadata about cache status.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#e6db74"&gt; &amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; start &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time()
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Check cache first&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; cached_answer, similarity &lt;span style="color:#f92672"&gt;=&lt;/span&gt; cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;get(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; cached_answer:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;answer&amp;#34;&lt;/span&gt;: cached_answer,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;from_cache&amp;#34;&lt;/span&gt;: &lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;similarity&amp;#34;&lt;/span&gt;: similarity,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;latency_ms&amp;#34;&lt;/span&gt;: (time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time() &lt;span style="color:#f92672"&gt;-&lt;/span&gt; start) &lt;span style="color:#f92672"&gt;*&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1000&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Cache miss — run full RAG&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; answer &lt;span style="color:#f92672"&gt;=&lt;/span&gt; rag_fn(query)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#75715e"&gt;# Store result in cache&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;set(query, answer)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; {
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;answer&amp;#34;&lt;/span&gt;: answer,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;from_cache&amp;#34;&lt;/span&gt;: &lt;span style="color:#66d9ef"&gt;False&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;similarity&amp;#34;&lt;/span&gt;: similarity,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;latency_ms&amp;#34;&lt;/span&gt;: (time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;time() &lt;span style="color:#f92672"&gt;-&lt;/span&gt; start) &lt;span style="color:#f92672"&gt;*&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1000&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; }
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Demo&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;dummy_rag&lt;/span&gt;(query: str) &lt;span style="color:#f92672"&gt;-&amp;gt;&lt;/span&gt; str:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;&amp;#34;&amp;#34;Simulates an expensive RAG call (sleep 2 seconds).&amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sleep(&lt;span style="color:#ae81ff"&gt;2&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# Simulate LLM latency&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; &lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;Answer to: &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;query&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cache &lt;span style="color:#f92672"&gt;=&lt;/span&gt; SemanticCache(similarity_threshold&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;0.92&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;queries &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;What is retrieval-augmented generation?&amp;#34;&lt;/span&gt;,
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;Explain RAG to me&amp;#34;&lt;/span&gt;, &lt;span style="color:#75715e"&gt;# Should hit cache&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;How does retrieval augmented generation work?&amp;#34;&lt;/span&gt;, &lt;span style="color:#75715e"&gt;# Should hit cache&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;What is a vector database?&amp;#34;&lt;/span&gt;, &lt;span style="color:#75715e"&gt;# Cache miss&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;Tell me about vector databases&amp;#34;&lt;/span&gt;, &lt;span style="color:#75715e"&gt;# Should hit cache&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;]
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; q &lt;span style="color:#f92672"&gt;in&lt;/span&gt; queries:
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; result &lt;span style="color:#f92672"&gt;=&lt;/span&gt; rag_with_cache(q, cache, dummy_rag)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; status &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;HIT&amp;#34;&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; result[&lt;span style="color:#e6db74"&gt;&amp;#34;from_cache&amp;#34;&lt;/span&gt;] &lt;span style="color:#66d9ef"&gt;else&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;MISS&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(&lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;[&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;status&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;] &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;q[:&lt;span style="color:#ae81ff"&gt;50&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt; — &lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;result[&lt;span style="color:#e6db74"&gt;&amp;#39;latency_ms&amp;#39;&lt;/span&gt;]&lt;span style="color:#e6db74"&gt;:&lt;/span&gt;&lt;span style="color:#e6db74"&gt;.0f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt;ms&amp;#34;&lt;/span&gt;)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;print(&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;\n&lt;/span&gt;&lt;span style="color:#e6db74"&gt;Cache stats:&amp;#34;&lt;/span&gt;, cache&lt;span style="color:#f92672"&gt;.&lt;/span&gt;stats())&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;hr&gt;
&lt;h2 id="production-redis--vector-cache"&gt;Production: Redis + Vector Cache&lt;a class="anchor" href="#production-redis--vector-cache"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;For production, store embeddings in Redis with vector search:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-4/vector-databases/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-4/vector-databases/</guid><description>&lt;h1 id="vector-databases"&gt;Vector Databases&lt;a class="anchor" href="#vector-databases"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;Embeddings without a vector database are like data without a database — unusable at scale.&lt;/p&gt;
&lt;/blockquote&gt;&lt;hr&gt;
&lt;h2 id="what-is-a-vector-database"&gt;What Is a Vector Database?&lt;a class="anchor" href="#what-is-a-vector-database"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A vector database stores high-dimensional vectors (embeddings) and answers the question: &lt;strong&gt;&amp;ldquo;Which vectors are most similar to this query vector?&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is the &amp;ldquo;Retrieve&amp;rdquo; in Retrieval-Augmented Generation.&lt;/p&gt;
&lt;pre class="mermaid"&gt;graph LR
 A[Query: &amp;#39;What is RAG?&amp;#39;] --&amp;gt; B[Embed Query]
 B --&amp;gt; C{Vector DB}
 C --&amp;gt; D[&amp;#34;[0.12, -0.34, 0.87, ...]&amp;#34;]
 D --&amp;gt; E[ANN Search: Top-K similar chunks]
 E --&amp;gt; F[Retrieved Context]&lt;/pre&gt;&lt;hr&gt;
&lt;h2 id="the-four-youll-use"&gt;The Four You&amp;rsquo;ll Use&lt;a class="anchor" href="#the-four-youll-use"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;DB&lt;/th&gt;
					&lt;th&gt;Type&lt;/th&gt;
					&lt;th&gt;Best For&lt;/th&gt;
					&lt;th&gt;Scale&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;FAISS&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Library&lt;/td&gt;
					&lt;td&gt;Research, GPU search&lt;/td&gt;
					&lt;td&gt;Millions (single machine)&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;ChromaDB&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Embedded DB&lt;/td&gt;
					&lt;td&gt;Prototyping, local dev&lt;/td&gt;
					&lt;td&gt;&amp;lt; 10M vectors&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Qdrant&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Production DB&lt;/td&gt;
					&lt;td&gt;Production RAG&lt;/td&gt;
					&lt;td&gt;Billions, distributed&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;PGVector&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Postgres ext.&lt;/td&gt;
					&lt;td&gt;Existing Postgres stack&lt;/td&gt;
					&lt;td&gt;Millions&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr&gt;
&lt;h2 id="faiss--facebook-ai-similarity-search"&gt;FAISS — Facebook AI Similarity Search&lt;a class="anchor" href="#faiss--facebook-ai-similarity-search"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;FAISS is a &lt;strong&gt;library&lt;/strong&gt;, not a server. It runs embedded in your Python process. Fastest for exact or approximate nearest-neighbor search.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/agent-evaluation-benchmarking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/agent-evaluation-benchmarking/</guid><description>&lt;h1 id="agent-benchmarking-and-evaluation"&gt;Agent Benchmarking and Evaluation&lt;a class="anchor" href="#agent-benchmarking-and-evaluation"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Benchmarking&lt;/strong&gt; compares agents on the same repeatable tasks. &lt;strong&gt;Evaluation&lt;/strong&gt; decides whether one is good enough for your real task.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Start with a benchmark when you want to explore a model or agent framework.
Then create a small evaluation set from your own work before trusting it with
users, tools, or data. An agent is not reliable because its answer sounds
convincing; evaluate the completed task, tool calls, artifacts, and failure
behaviour.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/agent-fundamentals/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/agent-fundamentals/</guid><description>&lt;h1 id="agent-fundamentals"&gt;Agent Fundamentals&lt;a class="anchor" href="#agent-fundamentals"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;An &lt;strong&gt;AI agent&lt;/strong&gt; is a model that can choose and use tools repeatedly to finish a goal.&lt;/p&gt;
&lt;/blockquote&gt;&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/ZvDkJsKE80k" title="AI Agents Explained — Tech With Tim"&gt;&lt;img src="https://img.youtube.com/vi/ZvDkJsKE80k/0.jpg" alt="AI Agents Explained — Tech With Tim (22 min, July 2026)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="agent-vs-chatbot"&gt;Agent vs chatbot&lt;a class="anchor" href="#agent-vs-chatbot"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Chatbot&lt;/th&gt;
					&lt;th&gt;Agent&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Usually produces one response&lt;/td&gt;
					&lt;td&gt;Can take several steps&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Mainly returns text&lt;/td&gt;
					&lt;td&gt;Can search, calculate, edit, or call APIs&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;User guides each turn&lt;/td&gt;
					&lt;td&gt;Agent chooses the next step&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Low autonomy&lt;/td&gt;
					&lt;td&gt;Controlled autonomy&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="the-agent-loop"&gt;The agent loop&lt;a class="anchor" href="#the-agent-loop"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 G[Goal] --&amp;gt; D[Decide next step]
 D --&amp;gt; T[Use a tool]
 T --&amp;gt; O[Observe result]
 O --&amp;gt; D
 D --&amp;gt;|Goal complete| F[Final answer]&lt;/pre&gt;&lt;p&gt;An agent repeats three simple actions:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/agent-memory-systems/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/agent-memory-systems/</guid><description>&lt;h1 id="agent-memory-systems"&gt;Agent Memory Systems&lt;a class="anchor" href="#agent-memory-systems"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Agent memory&lt;/strong&gt; is useful information saved now so the agent can use it later.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Saving an entire chat is history. Memory means selecting the small parts that will help a future task.&lt;/p&gt;
&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/BacJ6sEhqMo" title="The Four Types of Memory Every AI Agent Needs — IBM Technology"&gt;&lt;img src="https://img.youtube.com/vi/BacJ6sEhqMo/0.jpg" alt="The Four Types of Memory Every AI Agent Needs — IBM Technology (11 min, May 2026)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="types-of-memory"&gt;Types of memory&lt;a class="anchor" href="#types-of-memory"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Type&lt;/th&gt;
					&lt;th&gt;What it remembers&lt;/th&gt;
					&lt;th&gt;Example&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Working memory&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Current task and recent messages&lt;/td&gt;
					&lt;td&gt;Current order number&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Semantic memory&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Facts and preferences&lt;/td&gt;
					&lt;td&gt;User prefers Celsius&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Episodic memory&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Past events and outcomes&lt;/td&gt;
					&lt;td&gt;Last deployment failed&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Procedural memory&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;How to perform a task&lt;/td&gt;
					&lt;td&gt;Release checklist&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="simple-memory-flow"&gt;Simple memory flow&lt;a class="anchor" href="#simple-memory-flow"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 C[Conversation or event] --&amp;gt; S[Select useful fact]
 S --&amp;gt; V[Check permission and truth]
 V --&amp;gt; M[(Store memory)]
 M --&amp;gt; R[Retrieve when relevant]
 R --&amp;gt; A[Use in current task]&lt;/pre&gt;&lt;h2 id="where-memory-is-stored"&gt;Where memory is stored&lt;a class="anchor" href="#where-memory-is-stored"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Storage&lt;/th&gt;
					&lt;th&gt;Best use&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Conversation state&lt;/td&gt;
					&lt;td&gt;Current session&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;SQL/document database&lt;/td&gt;
					&lt;td&gt;Exact facts, profiles, and updates&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Vector database&lt;/td&gt;
					&lt;td&gt;Finding semantically similar memories&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Object storage&lt;/td&gt;
					&lt;td&gt;Large files, audio, images, and reports&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A normal database should usually be the source of truth. Add vector search only when fuzzy recall is useful.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/async-parallelism/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/async-parallelism/</guid><description>&lt;h1 id="async-and-parallelism-in-agent-systems"&gt;Async and Parallelism in Agent Systems&lt;a class="anchor" href="#async-and-parallelism-in-agent-systems"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Async&lt;/strong&gt; lets one program work on other tasks while it waits for a network, model, database, or file operation.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Agent systems spend a lot of time waiting, so async can reduce total waiting time.&lt;/p&gt;
&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/3E-ADzhr3W8" title="Asyncio in Python — Yash Jain"&gt;&lt;img src="https://img.youtube.com/vi/3E-ADzhr3W8/0.jpg" alt="Asyncio in Python — Yash Jain (15 min, November 2025)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="key-terms"&gt;Key terms&lt;a class="anchor" href="#key-terms"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Term&lt;/th&gt;
					&lt;th&gt;Simple meaning&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Synchronous&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Finish one task before starting the next.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Asynchronous&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Work on another task while waiting.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Concurrency&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Several tasks make progress during the same period.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Parallelism&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Several tasks run at the exact same time.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Coroutine&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Function declared with &lt;code&gt;async def&lt;/code&gt;.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;await&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Pause this coroutine until an operation finishes.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Event loop&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Schedules coroutines that are ready to run.&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="sequential-vs-concurrent"&gt;Sequential vs concurrent&lt;a class="anchor" href="#sequential-vs-concurrent"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 subgraph Sequential
 A1[Call A] --&amp;gt; B1[Call B] --&amp;gt; C1[Call C]
 end
 subgraph Concurrent
 S[Start] --&amp;gt; A2[Call A]
 S --&amp;gt; B2[Call B]
 S --&amp;gt; C2[Call C]
 end&lt;/pre&gt;&lt;h2 id="small-example"&gt;Small example&lt;a class="anchor" href="#small-example"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f92672"&gt;import&lt;/span&gt; asyncio
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;async&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;call_api&lt;/span&gt;(name):
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; asyncio&lt;span style="color:#f92672"&gt;.&lt;/span&gt;sleep(&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;) &lt;span style="color:#75715e"&gt;# represents network waiting&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;return&lt;/span&gt; &lt;span style="color:#e6db74"&gt;f&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;&lt;/span&gt;&lt;span style="color:#e6db74"&gt;{&lt;/span&gt;name&lt;span style="color:#e6db74"&gt;}&lt;/span&gt;&lt;span style="color:#e6db74"&gt; finished&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;async&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;def&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;main&lt;/span&gt;():
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; results &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; asyncio&lt;span style="color:#f92672"&gt;.&lt;/span&gt;gather(call_api(&lt;span style="color:#e6db74"&gt;&amp;#34;A&amp;#34;&lt;/span&gt;), call_api(&lt;span style="color:#e6db74"&gt;&amp;#34;B&amp;#34;&lt;/span&gt;))
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; print(results)
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;asyncio&lt;span style="color:#f92672"&gt;.&lt;/span&gt;run(main())&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both calls wait together, so this takes about one second instead of two.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/loop-engineering/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/loop-engineering/</guid><description>&lt;h1 id="loop-engineering"&gt;Loop Engineering&lt;a class="anchor" href="#loop-engineering"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Loop engineering&lt;/strong&gt; is the practice of designing how an AI agent repeatedly finds work, acts, checks the result, remembers progress, and stops safely.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;It is an emerging term from 2026. The underlying ideas—feedback loops, verification, state, and stop conditions—are established software and agent-design practices.&lt;/p&gt;
&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/4biXYSNkn9Y" title="Loop Engineering Explained — Caleb Writes Code"&gt;&lt;img src="https://img.youtube.com/vi/4biXYSNkn9Y/0.jpg" alt="Loop Engineering Explained — Caleb Writes Code (9 min, July 2026)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="the-loop"&gt;The loop&lt;a class="anchor" href="#the-loop"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 T[Trigger or new task] --&amp;gt; P[Pick next work item]
 P --&amp;gt; A[Agent acts]
 A --&amp;gt; V[Verify with evidence]
 V --&amp;gt;|Failed but recoverable| F[Give useful feedback]
 F --&amp;gt; A
 V --&amp;gt;|Passed| S[Save progress]
 S --&amp;gt; D{More work?}
 D --&amp;gt;|Yes| P
 D --&amp;gt;|No| H[Human review or stop]&lt;/pre&gt;&lt;h2 id="four-engineering-layers"&gt;Four engineering layers&lt;a class="anchor" href="#four-engineering-layers"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Layer&lt;/th&gt;
					&lt;th&gt;Main question&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Prompt engineering&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;What instruction should the model receive?&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Context engineering&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;What files, facts, and history should it see?&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Harness engineering&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Which tools, permissions, and checks support one run?&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Loop engineering&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;What triggers repeated runs, preserves progress, and stops them?&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;These layers work together. Loop engineering does not replace good prompts or good code.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/mcp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/mcp/</guid><description>&lt;h1 id="model-context-protocol-mcp"&gt;Model Context Protocol (MCP)&lt;a class="anchor" href="#model-context-protocol-mcp"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;MCP&lt;/strong&gt; is a standard way for AI applications to connect to external tools and data sources.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Instead of writing a different integration for every AI application, a developer can build one MCP server that compatible applications can use. MCP standardizes the connection; it does not replace access control, product design, or ordinary API design.&lt;/p&gt;
&lt;h2 id="videos"&gt;Videos&lt;a class="anchor" href="#videos"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/cGuyrANVi4A" title="How Model Context Protocol Actually Works — Google Cloud Tech"&gt;&lt;img src="https://img.youtube.com/vi/cGuyrANVi4A/0.jpg" alt="How Model Context Protocol Actually Works — Google Cloud Tech (8 min, June 2026)" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/multi-agent-systems/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/multi-agent-systems/</guid><description>&lt;h1 id="multi-agent-systems"&gt;Multi-Agent Systems&lt;a class="anchor" href="#multi-agent-systems"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;A &lt;strong&gt;multi-agent system&lt;/strong&gt; uses several agents with different roles to complete one larger goal.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;More agents do not automatically mean better results. Use them only when tasks can be separated or require different tools and permissions.&lt;/p&gt;
&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/sWH0T4Zez6I" title="Multi-Agent Systems Explained — IBM Technology"&gt;&lt;img src="https://img.youtube.com/vi/sWH0T4Zez6I/0.jpg" alt="Multi-Agent Systems Explained — IBM Technology (8 min, December 2025)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="supervisor-pattern"&gt;Supervisor pattern&lt;a class="anchor" href="#supervisor-pattern"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart TD
 U[User goal] --&amp;gt; S[Supervisor]
 S --&amp;gt; R[Research agent]
 S --&amp;gt; C[Coding agent]
 R --&amp;gt; S
 C --&amp;gt; S
 S --&amp;gt; V[Check and combine]
 V --&amp;gt; U&lt;/pre&gt;&lt;p&gt;The supervisor divides the work, gives each worker a clear task, and combines the results.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/sandboxing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/sandboxing/</guid><description>&lt;h1 id="sandboxing-agent-code"&gt;Sandboxing Agent Code&lt;a class="anchor" href="#sandboxing-agent-code"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;A &lt;strong&gt;sandbox&lt;/strong&gt; is an isolated environment that limits what agent-generated code can access and damage.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Sandboxing is a security boundary, not a particular product. Containers,
application kernels, microVMs, virtual machines, and managed code-execution
services can all be useful. The right choice depends on what the code can do,
which data it can reach, and the cost of a breakout or mistake.&lt;/p&gt;
&lt;h2 id="videos"&gt;Videos&lt;a class="anchor" href="#videos"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/bm6jegefGyY" title="Create a Python Sandbox for Agents — Trelis Research"&gt;&lt;img src="https://img.youtube.com/vi/bm6jegefGyY/0.jpg" alt="Create a Python Sandbox for Agents — Trelis Research (16 min)" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/specialized-agents/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/specialized-agents/</guid><description>&lt;h1 id="specialized-agents"&gt;Specialized Agents&lt;a class="anchor" href="#specialized-agents"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;A &lt;strong&gt;specialized agent&lt;/strong&gt; is designed for one type of work and receives only the tools, data, and permissions needed for that work.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;Specialization is more than giving the model a role in a prompt. The tools and checks must also match the role.&lt;/p&gt;
&lt;h2 id="short-video"&gt;Short video&lt;a class="anchor" href="#short-video"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/fXizBc03D7E" title="Five Types of AI Agents — IBM Technology"&gt;&lt;img src="https://img.youtube.com/vi/fXizBc03D7E/0.jpg" alt="Five Types of AI Agents — IBM Technology (10 min, April 2025)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="basic-design"&gt;Basic design&lt;a class="anchor" href="#basic-design"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;pre class="mermaid"&gt;flowchart LR
 G[Clear goal] --&amp;gt; A[Specialized agent]
 D[Relevant data] --&amp;gt; A
 A --&amp;gt; T[Limited tools]
 T --&amp;gt; R[Result]
 R --&amp;gt; V[Domain-specific check]&lt;/pre&gt;&lt;h2 id="useful-specializations"&gt;Useful specializations&lt;a class="anchor" href="#useful-specializations"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Agent&lt;/th&gt;
					&lt;th&gt;Tools&lt;/th&gt;
					&lt;th&gt;How to check it&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Browser agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Browser automation&lt;/td&gt;
					&lt;td&gt;Page state or screenshot&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Coding agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Editor, shell, tests&lt;/td&gt;
					&lt;td&gt;Tests, types, and diff review&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Data agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;SQL and notebooks&lt;/td&gt;
					&lt;td&gt;Constraints and totals&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Research agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Search and document tools&lt;/td&gt;
					&lt;td&gt;Primary-source citations&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Document agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;PDF/Office parsers&lt;/td&gt;
					&lt;td&gt;Rendered-page review&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Operations agent&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Logs and runbooks&lt;/td&gt;
					&lt;td&gt;Health check and rollback state&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="browser-agents"&gt;Browser agents&lt;a class="anchor" href="#browser-agents"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Browser agents can read pages, fill forms, and click controls. Prefer page roles and labels such as “Search” or “Submit” over screen coordinates because they are more stable.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-5/tool-calling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-5/tool-calling/</guid><description>&lt;h1 id="function--tool-calling"&gt;Function / Tool Calling&lt;a class="anchor" href="#function--tool-calling"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Tool calling&lt;/strong&gt; lets a model ask an application to run a function.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;The model does not execute the function. It returns the tool name and arguments; the application validates and runs them.&lt;/p&gt;
&lt;h2 id="videos"&gt;Videos&lt;a class="anchor" href="#videos"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/h8gMhXYAv1k" title="What is Tool Calling? — IBM Technology"&gt;&lt;img src="https://img.youtube.com/vi/h8gMhXYAv1k/0.jpg" alt="What is Tool Calling? — IBM Technology (5 min, January 2025)" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://youtu.be/gMeTK6zzaO4" title="LLM Function Calling — AI Tools Deep Dive"&gt;&lt;img src="https://img.youtube.com/vi/gMeTK6zzaO4/0.jpg" alt="LLM Function Calling — AI Tools Deep Dive (31 min)" /&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/anti-bot-patterns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/anti-bot-patterns/</guid><description>&lt;h1 id="anti-bot-patterns"&gt;Anti-bot Patterns&lt;a class="anchor" href="#anti-bot-patterns"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Most blocks aren&amp;rsquo;t Cloudflare-grade. Learn the short ladder of defences sites use — and the honest response to each rung.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~12 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Before you reach for stealth browsers, know the ladder. Nine times out of ten a &amp;ldquo;block&amp;rdquo; is something simple — a missing header, too many requests, or a session cookie you didn&amp;rsquo;t carry. Each rung has a &lt;em&gt;correct&lt;/em&gt; response and an &lt;em&gt;arms-race&lt;/em&gt; response; prefer the correct one.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/authenticated-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/authenticated-scraping/</guid><description>&lt;h1 id="authenticated-scraping"&gt;Authenticated Scraping&lt;a class="anchor" href="#authenticated-scraping"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Carry a login session in code — CSRF tokens, cookies, saved browser state — and recognise when the login &lt;em&gt;is&lt;/em&gt; the line you shouldn&amp;rsquo;t cross.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-6/playwright-selenium/"&gt;Playwright &amp;amp; Selenium&lt;/a&gt; · &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Plenty of data sits behind a login. Sometimes you&amp;rsquo;re clearly entitled to it — your own account, your own app, or an API that issues you a token. Sometimes the login is precisely the access control you must not defeat.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/change-detection-dedup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/change-detection-dedup/</guid><description>&lt;h1 id="change-detection--dedup"&gt;Change Detection &amp;amp; Dedup&lt;a class="anchor" href="#change-detection--dedup"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;A scheduled scraper that re-saves the same rows every night is just an expensive clock. Store state, hash content, and record only what actually changed.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/duckdb-parquet/"&gt;DuckDB + Parquet&lt;/a&gt; · &lt;a href="/2026-05/week-6/scheduled-scraping/"&gt;Scheduled Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The first run of a scraper is the easy one. The interesting question is the second: &lt;em&gt;what&amp;rsquo;s new?&lt;/em&gt; Answer it with two cheap ideas — a &lt;strong&gt;stable ID&lt;/strong&gt; for every record, and a &lt;strong&gt;content hash&lt;/strong&gt; to detect edits.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/cloudflare-bot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/cloudflare-bot/</guid><description>&lt;h1 id="cloudflare-bot-protection"&gt;Cloudflare Bot Protection&lt;a class="anchor" href="#cloudflare-bot-protection"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Understand &lt;em&gt;exactly&lt;/em&gt; how Cloudflare decides you&amp;rsquo;re a bot — then use the legitimate ways through, and know what it costs to fetch a page anyway.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/anti-bot-patterns/"&gt;Anti-bot Patterns&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Cloudflare sits in front of a large share of the web. A plain &lt;code&gt;httpx&lt;/code&gt; or &lt;code&gt;requests&lt;/code&gt; call often gets a &lt;code&gt;403&lt;/code&gt; or an endless &amp;ldquo;checking your browser&amp;rdquo; loop, while Chrome loads the same page instantly. That gap isn&amp;rsquo;t magic — it&amp;rsquo;s four specific signals. Once you can name them, you know your options.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/document-parsing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/document-parsing/</guid><description>&lt;h1 id="document-parsing"&gt;Document Parsing&lt;a class="anchor" href="#document-parsing"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Most of the world&amp;rsquo;s data is trapped in PDFs, Word files, and scans. Getting it out cleanly — with tables intact — is its own skill.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/html-to-markdown/"&gt;HTML → Markdown&lt;/a&gt; · &lt;a href="/2026-05/week-6/vision-models-for-scraping/"&gt;Vision Models for Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A PDF isn&amp;rsquo;t a document format so much as a set of drawing instructions. There&amp;rsquo;s often no &amp;ldquo;table&amp;rdquo; in there at all — just text positioned at coordinates that &lt;em&gt;look&lt;/em&gt; like a table. That&amp;rsquo;s why naive extraction produces scrambled columns, and why picking the right tool matters.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/duckdb-parquet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/duckdb-parquet/</guid><description>&lt;h1 id="duckdb--parquet-and-sqlite-for-state"&gt;DuckDB + Parquet (and SQLite for state)&lt;a class="anchor" href="#duckdb--parquet-and-sqlite-for-state"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Two databases, two jobs: SQLite remembers what your scraper has seen; DuckDB answers questions about what it collected — straight off Parquet files, no server.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-1/05-sqlite/"&gt;SQLite&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Scraped data has two very different access patterns, and using one tool for both is why people end up with a 4 GB CSV they can&amp;rsquo;t open.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/google-dork/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/google-dork/</guid><description>&lt;h1 id="google-dorking"&gt;Google Dorking&lt;a class="anchor" href="#google-dorking"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Turn a search box into a precision data-sourcing tool — find the exact files and datasets you need, then automate it into a reproducible pipeline.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Dorking&amp;rdquo; is just using search operators well. The same operators that find a public dataset in seconds also reveal what an organisation has &lt;em&gt;accidentally&lt;/em&gt; left indexed — so this is both a sourcing skill and the first move in a footprint audit. Here we focus on &lt;strong&gt;sourcing + automation&lt;/strong&gt;; the defensive exposure-hunting side is &lt;a href="/2026-05/week-7/dorking-recon/"&gt;Week 7 → Dorking for Recon&lt;/a&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/hidden-json-apis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/hidden-json-apis/</guid><description>&lt;h1 id="hidden-json-apis"&gt;Hidden JSON APIs&lt;a class="anchor" href="#hidden-json-apis"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The page you want to scrape has already done the work for you.&lt;/strong&gt; Find the JSON endpoint its own JavaScript calls, and skip the HTML entirely.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-1/06-http-clients/"&gt;HTTP clients&lt;/a&gt; · browser DevTools (Network tab)&lt;/p&gt;
&lt;p&gt;Most modern sites load a nearly-empty HTML shell, then fetch their real data as JSON from a backend API. If you find that request, you get clean, structured data — no brittle CSS selectors, no headless browser, often 100× faster. This is the &lt;strong&gt;first thing to try&lt;/strong&gt; on any dynamic site, before you reach for Playwright.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/html-to-markdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/html-to-markdown/</guid><description>&lt;h1 id="html--markdown-for-llms"&gt;HTML → Markdown for LLMs&lt;a class="anchor" href="#html--markdown-for-llms"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Raw HTML is 90% navigation, scripts, and cookie banners. Strip it to clean Markdown and you cut your token bill while improving the model&amp;rsquo;s answers.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~12 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-6/document-parsing/"&gt;Document Parsing&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Feeding raw HTML to an LLM wastes tokens on markup the model doesn&amp;rsquo;t need and buries the actual content in boilerplate. Converting to Markdown first is one of the highest-leverage steps in any scrape-to-LLM pipeline.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/image-processing-pipeline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/image-processing-pipeline/</guid><description>&lt;h1 id="image-processing-pipeline"&gt;Image Processing Pipeline&lt;a class="anchor" href="#image-processing-pipeline"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Scraped images are rarely usable as-is. Deduplicate, normalise, and crop them before they reach a model or a database.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~7 min read · ~12 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/vision-models-for-scraping/"&gt;Vision Models for Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Scrape a few thousand images and you&amp;rsquo;ll have duplicates at different resolutions, EXIF-rotated photos that appear sideways, and 8 MB PNGs where a 200 KB JPEG would do. Fix that in a pipeline, once.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/legal-ethical-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/legal-ethical-scraping/</guid><description>&lt;h1 id="legal--ethical-scraping"&gt;Legal &amp;amp; Ethical Scraping&lt;a class="anchor" href="#legal--ethical-scraping"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;It&amp;rsquo;s public&amp;rdquo; does not mean &amp;ldquo;you&amp;rsquo;re allowed.&amp;rdquo;&lt;/strong&gt; Learn to answer &lt;em&gt;&amp;ldquo;can I scrape this?&amp;rdquo;&lt;/em&gt; before you write a line of code — and how to stay on the right side of the line.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: nothing · read this &lt;strong&gt;before&lt;/strong&gt; every other Week 6 page&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;ℹ️ This is a practical engineering guide, &lt;strong&gt;not legal advice&lt;/strong&gt;. Laws differ by country and change often. When money, personal data, or a company&amp;rsquo;s core assets are involved, ask a lawyer.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/osint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/osint/</guid><description>&lt;h1 id="osint--infrastructure--public-records"&gt;OSINT — Infrastructure &amp;amp; Public Records&lt;a class="anchor" href="#osint--infrastructure--public-records"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Build a picture of an organisation&amp;rsquo;s public footprint — domains, certificates, code, filings — from open sources, and know exactly how far a single source can be trusted.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/wayback-commoncrawl/"&gt;Wayback &amp;amp; Common Crawl&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Open-source intelligence (OSINT) is assembling a verified picture from public sources. This page scopes to &lt;strong&gt;infrastructure and public records&lt;/strong&gt; — the material that matters for data science and security analysis. Profiling people is a different discipline with a much heavier duty of care; that lives in &lt;a href="/2026-05/week-7/person-social-osint/"&gt;Week 7 → Person &amp;amp; Social OSINT&lt;/a&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/pagination-infinite-scroll/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/pagination-infinite-scroll/</guid><description>&lt;h1 id="pagination--infinite-scroll"&gt;Pagination &amp;amp; Infinite Scroll&lt;a class="anchor" href="#pagination--infinite-scroll"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Load more&amp;rdquo; is a button wired to a loop.&lt;/strong&gt; Find what it increments — a page number or a cursor token — and the signal that says stop, and you can fetch every result without clicking it.&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~6 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-1/06-http-clients/"&gt;HTTP clients&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Any list longer than one screen — search results, a catalog, a feed — is paginated somehow: page numbers, a &amp;ldquo;Load more&amp;rdquo; button, or an endless scrollbar. Reach for this the moment you see one; skip it if the first response already returns everything (check for a &lt;code&gt;total&lt;/code&gt;/&lt;code&gt;count&lt;/code&gt; field first).&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/playwright-advanced/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/playwright-advanced/</guid><description>&lt;h1 id="playwright-advanced"&gt;Playwright Advanced&lt;a class="anchor" href="#playwright-advanced"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Intercept the network, reuse a saved login, and block the junk — the difference between a browser script that crawls and one that flies.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/playwright-selenium/"&gt;Playwright &amp;amp; Selenium&lt;/a&gt; · &lt;a href="/2026-05/week-6/authenticated-scraping/"&gt;Authenticated Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Once basic automation works, three techniques make it production-grade: &lt;strong&gt;request interception&lt;/strong&gt;, &lt;strong&gt;saved authentication state&lt;/strong&gt;, and &lt;strong&gt;tracing&lt;/strong&gt;. Together they typically cut runtime by 3–5× and eliminate most flakiness.&lt;/p&gt;
&lt;h2 id="try-it-in-5-minutes--make-it-5-faster"&gt;Try it in 5 minutes — make it 5× faster&lt;a class="anchor" href="#try-it-in-5-minutes--make-it-5-faster"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Images, fonts, ads, and analytics are pure overhead when you only want text. Block them at the network layer:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/playwright-selenium/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/playwright-selenium/</guid><description>&lt;h1 id="playwright--selenium"&gt;Playwright &amp;amp; Selenium&lt;a class="anchor" href="#playwright--selenium"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;When there&amp;rsquo;s genuinely no API behind the page, drive a real browser — and drive it so it waits for content instead of guessing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt; · &lt;a href="/2026-05/week-1/06-http-clients/"&gt;HTTP clients&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Browser automation is the heavyweight option: it renders JavaScript, executes the page&amp;rsquo;s own code, and sees exactly what a user sees. It&amp;rsquo;s also 10–100× slower than an HTTP request and far more fragile. Use it &lt;strong&gt;after&lt;/strong&gt; you&amp;rsquo;ve checked for a &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;hidden JSON API&lt;/a&gt;, not before.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/rate-limits-retries-caching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/rate-limits-retries-caching/</guid><description>&lt;h1 id="rate-limits-retries--caching"&gt;Rate Limits, Retries &amp;amp; Caching&lt;a class="anchor" href="#rate-limits-retries--caching"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Scrape so politely you never get banned, and so efficiently you never fetch the same page twice.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-1/06-http-clients/"&gt;HTTP clients&lt;/a&gt; · &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The difference between a scraper that runs for months and one that&amp;rsquo;s blocked on day one is rarely cleverness — it&amp;rsquo;s restraint. Three habits do almost all the work: &lt;strong&gt;cap your concurrency&lt;/strong&gt;, &lt;strong&gt;back off when told to&lt;/strong&gt;, and &lt;strong&gt;cache everything&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/scheduled-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/scheduled-scraping/</guid><description>&lt;h1 id="scheduled-scraping"&gt;Scheduled Scraping&lt;a class="anchor" href="#scheduled-scraping"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Data is only useful if it&amp;rsquo;s fresh. Put your scraper on a free cron, make it idempotent, and let it build a time-series while you sleep.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/change-detection-dedup/"&gt;Change Detection &amp;amp; Dedup&lt;/a&gt; · &lt;a href="/2026-05/week-1/04-git-github/"&gt;GitHub Actions&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A one-off scrape is a snapshot. A &lt;em&gt;scheduled&lt;/em&gt; scrape is a dataset that gets more valuable every day — and GitHub Actions will run it for free.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/sitemaps-rss-jsonld/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/sitemaps-rss-jsonld/</guid><description>&lt;h1 id="sitemaps-rss--structured-data"&gt;Sitemaps, RSS &amp;amp; Structured Data&lt;a class="anchor" href="#sitemaps-rss--structured-data"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Before you write a single selector, check whether the site is already handing out its data — as a URL index, a change feed, or clean JSON embedded in the page.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Sites publish structured data on purpose — for search engines, feed readers, and social previews. It&amp;rsquo;s stable, it&amp;rsquo;s meant to be machine-read, and almost nobody checks for it before writing a fragile HTML parser.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/speech-ai/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/speech-ai/</guid><description>&lt;h1 id="speech-ai"&gt;Speech AI&lt;a class="anchor" href="#speech-ai"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Audio is a data source you can query. Transcribe it, timestamp it, and it becomes searchable text like anything else you scraped.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~12 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/video-understanding/"&gt;Video Understanding&lt;/a&gt; · &lt;a href="/2026-05/week-2/11-local-llms-1-basics/"&gt;Local LLMs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Podcasts, lectures, earnings calls, support recordings — enormous amounts of information exist only as speech. Speech-to-text (STT) turns it into text you can search, chunk, and feed to an LLM.&lt;/p&gt;
&lt;h2 id="try-it-in-5-minutes--transcribe-with-timestamps"&gt;Try it in 5 minutes — transcribe with timestamps&lt;a class="anchor" href="#try-it-in-5-minutes--transcribe-with-timestamps"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;faster-whisper&lt;/code&gt; runs Whisper on a CTranslate2 backend — several times quicker than the reference implementation and comfortable on CPU with a small model:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/video-understanding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/video-understanding/</guid><description>&lt;h1 id="video-understanding"&gt;Video Understanding&lt;a class="anchor" href="#video-understanding"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;A video is frames plus audio plus time. Split it into those three, and a problem that looked impossible becomes three you already know how to solve.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~12 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/speech-ai/"&gt;Speech AI&lt;/a&gt; · &lt;a href="/2026-05/week-6/vision-models-for-scraping/"&gt;Vision Models for Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Never treat video as an opaque blob. Decompose it: &lt;strong&gt;audio&lt;/strong&gt; → &lt;a href="/2026-05/week-6/speech-ai/"&gt;transcript&lt;/a&gt;; &lt;strong&gt;frames&lt;/strong&gt; → &lt;a href="/2026-05/week-6/vision-models-for-scraping/"&gt;vision models&lt;/a&gt;; &lt;strong&gt;time&lt;/strong&gt; → the index that ties them together.&lt;/p&gt;
&lt;h2 id="try-it-in-5-minutes--ffmpeg-is-the-whole-toolkit"&gt;Try it in 5 minutes — ffmpeg is the whole toolkit&lt;a class="anchor" href="#try-it-in-5-minutes--ffmpeg-is-the-whole-toolkit"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# One frame per second, numbered&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i video.mp4 -vf fps&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#ae81ff"&gt;1&lt;/span&gt; frame_%04d.jpg
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Audio only, for transcription&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -i video.mp4 -vn -acodec libmp3lame audio.mp3
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# A single frame at 01:23&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffmpeg -ss 00:01:23 -i video.mp4 -frames:v &lt;span style="color:#ae81ff"&gt;1&lt;/span&gt; shot.jpg
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# Duration and stream info as JSON&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ffprobe -v quiet -print_format json -show_format -show_streams video.mp4&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;✅ Extraction, sampling, and metadata — four commands cover most of what a video pipeline needs.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/vision-models-for-scraping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/vision-models-for-scraping/</guid><description>&lt;h1 id="vision-models-for-scraping"&gt;Vision Models for Scraping&lt;a class="anchor" href="#vision-models-for-scraping"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;When the data is a chart, a scanned table, or a UI built to defeat parsing — screenshot it and ask a vision model for JSON.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/playwright-selenium/"&gt;Playwright &amp;amp; Selenium&lt;/a&gt; · &lt;a href="/2026-05/week-3/structured-output/"&gt;Structured Output&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Some data simply isn&amp;rsquo;t in the DOM: values baked into an image, a canvas-rendered chart, a scanned PDF page, or a deliberately obfuscated layout. A vision-language model (VLM) reads the &lt;em&gt;rendered pixels&lt;/em&gt; the way a person would.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-6/wayback-commoncrawl/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-6/wayback-commoncrawl/</guid><description>&lt;h1 id="wayback-machine--common-crawl"&gt;Wayback Machine &amp;amp; Common Crawl&lt;a class="anchor" href="#wayback-machine--common-crawl"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Someone already crawled the web for you. Get the page — and its entire history — without sending the target a single request.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~15 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt; · &lt;a href="/2026-05/week-6/hidden-json-apis/"&gt;Hidden JSON APIs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Two public archives cover a large share of the web: the &lt;strong&gt;Internet Archive&amp;rsquo;s Wayback Machine&lt;/strong&gt; (snapshots of individual URLs over time) and &lt;strong&gt;Common Crawl&lt;/strong&gt; (petabyte-scale crawls released free for research). Reach for them when a site blocks you, when you need &lt;em&gt;history&lt;/em&gt; rather than the current page, or when you need breadth no polite scraper could achieve.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/01-github-actions-advanced/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/01-github-actions-advanced/</guid><description>&lt;h1 id="github-actions--advanced"&gt;GitHub Actions — Advanced&lt;a class="anchor" href="#github-actions--advanced"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Once a workflow works, make it fast, safe, and reusable: cache the slow parts, matrix the repetitive parts, and stop handing every job write access to your repo.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-1/04-git-github/"&gt;Git &amp;amp; GitHub&lt;/a&gt; · &lt;a href="/2026-05/week-6/scheduled-scraping/"&gt;Scheduled Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ve used Actions to run a scheduled scraper. Production CI/CD adds three demands: it must be &lt;strong&gt;fast&lt;/strong&gt; (nobody waits 20 minutes), &lt;strong&gt;safe&lt;/strong&gt; (a workflow is code with access to your secrets), and &lt;strong&gt;reusable&lt;/strong&gt; (don&amp;rsquo;t copy-paste YAML across ten repos).&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/02-advanced-docker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/02-advanced-docker/</guid><description>&lt;h1 id="advanced-docker"&gt;Advanced Docker&lt;a class="anchor" href="#advanced-docker"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;A 1.2 GB image that rebuilds from scratch on every code change is a build-system bug. Multi-stage builds and correct layer order fix both size and speed.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-2/06-docker-compose/"&gt;Docker &amp;amp; Compose&lt;/a&gt; · &lt;a href="/2026-05/week-7/01-github-actions-advanced/"&gt;GitHub Actions Advanced&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You can already write a Dockerfile. Production adds three requirements: &lt;strong&gt;small&lt;/strong&gt; (fast pulls, less to attack), &lt;strong&gt;cached&lt;/strong&gt; (rebuild in seconds), and &lt;strong&gt;safe&lt;/strong&gt; (no root, no secrets baked in).&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/03-llm-security-offensive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/03-llm-security-offensive/</guid><description>&lt;h1 id="llm-security--offensive"&gt;LLM Security — Offensive&lt;a class="anchor" href="#llm-security--offensive"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Attack your own LLM app before a stranger does. Prompt injection is not a hypothetical — it&amp;rsquo;s the default behaviour of a system that can&amp;rsquo;t tell instructions from data.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/05-owasp-llm-top-10/"&gt;OWASP LLM Top 10&lt;/a&gt; · &lt;a href="/2026-05/week-3/01-prompt-engineering-1-foundations/"&gt;Prompt Engineering&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;An LLM receives one flat stream of text. Your careful system prompt and a hostile sentence inside a scraped PDF arrive in the &lt;em&gt;same channel&lt;/em&gt;. That&amp;rsquo;s the root cause of nearly every LLM attack — and why &amp;ldquo;just tell it to ignore malicious instructions&amp;rdquo; doesn&amp;rsquo;t work.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/04-llm-safety-defensive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/04-llm-safety-defensive/</guid><description>&lt;h1 id="llm-safety--defensive"&gt;LLM Safety — Defensive&lt;a class="anchor" href="#llm-safety--defensive"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;You cannot make a model immune to prompt injection. You can make a successful injection worthless — by limiting what the system is able to do.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/03-llm-security-offensive/"&gt;LLM Security — Offensive&lt;/a&gt; · &lt;a href="/2026-05/week-5/sandboxing/"&gt;Sandboxing Agent Code&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The defensive mindset is &lt;strong&gt;assume the model will be compromised&lt;/strong&gt;. Design so that a model doing the worst possible thing still can&amp;rsquo;t cause serious harm. Filtering inputs helps at the margin; architecture is what actually saves you.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/05-owasp-llm-top-10/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/05-owasp-llm-top-10/</guid><description>&lt;h1 id="owasp-top-10-for-llm-applications"&gt;OWASP Top 10 for LLM Applications&lt;a class="anchor" href="#owasp-top-10-for-llm-applications"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The industry&amp;rsquo;s shared checklist of how LLM applications actually get broken — and the one to run your own project against before it ships.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/03-llm-security-offensive/"&gt;LLM Security — Offensive&lt;/a&gt; · &lt;a href="/2026-05/week-7/04-llm-safety-defensive/"&gt;LLM Safety — Defensive&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/"&gt;OWASP&lt;/a&gt; maintains the reference list of LLM-specific risks. It&amp;rsquo;s the vocabulary security teams use, so knowing the codes is genuinely useful in a review: &amp;ldquo;that&amp;rsquo;s LLM06&amp;rdquo; lands faster than a paragraph.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/06-vms-ssh/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/06-vms-ssh/</guid><description>&lt;h1 id="vms--ssh"&gt;VMs &amp;amp; SSH&lt;a class="anchor" href="#vms--ssh"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Sometimes you just need a machine. Rent one, reach it safely with keys, and keep your work alive after you disconnect.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-1/03-bash-scripting/"&gt;Bash Scripting&lt;/a&gt; · &lt;a href="/2026-05/week-2/07-deployment-platforms/"&gt;Deployment Platforms&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Serverless covers most workloads, but some jobs want a persistent box: a long scrape, a GPU fine-tune, a database you control. That means a VM — and SSH is how you live on it.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/07-serverless-functions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/07-serverless-functions/</guid><description>&lt;h1 id="serverless-functions"&gt;Serverless Functions&lt;a class="anchor" href="#serverless-functions"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Deploy code without owning a server, pay only while it runs, and scale to zero when nobody&amp;rsquo;s looking — provided you design for a process that can vanish at any moment.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-2/07-deployment-platforms/"&gt;Deployment Platforms&lt;/a&gt; · &lt;a href="/2026-05/week-7/02-advanced-docker/"&gt;Advanced Docker&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Serverless is the natural home for the workloads this course produces: a scheduled scraper, a webhook receiver, an inference endpoint used a few hundred times a day. You ship a function or container; the platform handles machines, scaling, and idle cost.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/08-terraform-iac/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/08-terraform-iac/</guid><description>&lt;h1 id="terraform--infrastructure-as-code"&gt;Terraform &amp;amp; Infrastructure as Code&lt;a class="anchor" href="#terraform--infrastructure-as-code"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Clicking through a cloud console is undocumented, unrepeatable, and unreviewable. Declare your infrastructure in files, and &lt;code&gt;git log&lt;/code&gt; becomes your change history.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/07-serverless-functions/"&gt;Serverless Functions&lt;/a&gt; · &lt;a href="/2026-05/week-1/04-git-github/"&gt;Git &amp;amp; GitHub&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Infrastructure as Code (IaC) means your cloud resources are declared in version-controlled files. You describe the &lt;strong&gt;desired end state&lt;/strong&gt;; Terraform computes the diff and applies it. The wins are reproducibility, code review for infrastructure, and one command to tear a whole environment down.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/09-cost-alerting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/09-cost-alerting/</guid><description>&lt;h1 id="cost-alerting--budgets"&gt;Cost Alerting &amp;amp; Budgets&lt;a class="anchor" href="#cost-alerting--budgets"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The cloud bill nobody checks is the one that ruins a month. Cap what you can, alert on the rest, and never let an agent loop spend money unattended.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~8 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/07-serverless-functions/"&gt;Serverless Functions&lt;/a&gt; · &lt;a href="/2026-05/week-5/loop-engineering/"&gt;Loop Engineering&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Students in this course run LLM APIs, autoscaling services, and agent loops — three of the best ways ever invented to spend money by accident. A retry loop against a paid API can burn a semester&amp;rsquo;s budget overnight.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/10-pubsub-event-driven/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/10-pubsub-event-driven/</guid><description>&lt;h1 id="pubsub--event-driven-architecture"&gt;Pub/Sub &amp;amp; Event-Driven Architecture&lt;a class="anchor" href="#pubsub--event-driven-architecture"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Stop making the caller wait. Drop a message on a queue, return immediately, and let workers do the slow part — with retries you didn&amp;rsquo;t have to write.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-7/07-serverless-functions/"&gt;Serverless Functions&lt;/a&gt; · &lt;a href="/2026-05/week-5/async-parallelism/"&gt;Async &amp;amp; Parallelism&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Your scraper takes 90 seconds. Your &lt;a href="/2026-05/week-7/07-serverless-functions/"&gt;serverless request&lt;/a&gt; times out at 60. The fix isn&amp;rsquo;t a bigger timeout — it&amp;rsquo;s splitting the work: accept the job, queue it, respond instantly, and process it elsewhere.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/cloudflare-defender/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/cloudflare-defender/</guid><description>&lt;h1 id="cloudflare--the-defenders-side"&gt;Cloudflare — The Defender&amp;rsquo;s Side&lt;a class="anchor" href="#cloudflare--the-defenders-side"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;You&amp;rsquo;ve spent Week 6 getting past bot protection. Now put it in front of your own API and watch the traffic from the other side of the glass.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/cloudflare-bot/"&gt;Cloudflare Bot Protection&lt;/a&gt; · &lt;a href="/2026-05/week-2/10-cloudflare-tunnels/"&gt;Cloudflare Tunnels&lt;/a&gt; · &lt;a href="/2026-05/week-2/01-fastapi/"&gt;FastAPI&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Every API you ship will be scraped, credential-stuffed, and scanned. The defences you studied as obstacles in Week 6 are the ones you now have to &lt;em&gt;configure&lt;/em&gt; — and the interesting part is that the goal is never &amp;ldquo;block all bots.&amp;rdquo; It&amp;rsquo;s to let the right ones through cheaply while making the wrong ones expensive.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/dorking-recon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/dorking-recon/</guid><description>&lt;h1 id="dorking-for-recon--exposure"&gt;Dorking for Recon &amp;amp; Exposure&lt;a class="anchor" href="#dorking-for-recon--exposure"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Search engines have already indexed your mistakes. Find them before someone else does — on domains you own.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~9 min read · ~20 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/google-dork/"&gt;Google Dorking&lt;/a&gt; · &lt;a href="/2026-05/week-7/05-owasp-llm-top-10/"&gt;OWASP LLM Top 10&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="/2026-05/week-6/google-dork/"&gt;Week 6&lt;/a&gt; taught operators as a &lt;em&gt;sourcing&lt;/em&gt; skill. The same operators are the first tool in an attacker&amp;rsquo;s kit — and therefore the first in a defender&amp;rsquo;s. This page is about &lt;strong&gt;auditing your own external footprint&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-7/person-social-osint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-7/person-social-osint/</guid><description>&lt;h1 id="person--social-osint"&gt;Person &amp;amp; Social OSINT&lt;a class="anchor" href="#person--social-osint"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The same techniques that profile a person will show you what the internet already knows about &lt;em&gt;you&lt;/em&gt;. Learn both halves — and practise only on yourself.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~25 min hands-on
🔗 needs: &lt;a href="/2026-05/week-6/osint/"&gt;OSINT — Infrastructure &amp;amp; Records&lt;/a&gt; · &lt;a href="/2026-05/week-6/legal-ethical-scraping/"&gt;Legal &amp;amp; Ethical Scraping&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="/2026-05/week-6/osint/"&gt;Week 6&lt;/a&gt; scoped OSINT to domains, certificates, and corporate records. This page covers the part aimed at &lt;strong&gt;people&lt;/strong&gt; — because you cannot defend against a technique you don&amp;rsquo;t understand, and because as a data scientist you will be handed personal datasets and asked what&amp;rsquo;s safe to publish.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/01-cloud-storage-ml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/01-cloud-storage-ml/</guid><description>&lt;h1 id="cloud-storage-for-ml"&gt;Cloud Storage for ML&lt;a class="anchor" href="#cloud-storage-for-ml"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=F4XFrHhhLow"&gt;&lt;img src="https://img.youtube.com/vi/F4XFrHhhLow/0.jpg" alt="Create and use a Cloud Storage bucket" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;A model is not a single file. It is a chain of data, code, weights, metrics, and predictions. Cloud storage gives that chain a home that is not your laptop.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~10 min read · ~25 min hands-on
🔗 needs: &lt;a href="/2026-05/week-2/05-config-management/"&gt;Config Management&lt;/a&gt; · &lt;a href="/2026-05/week-7/09-cost-alerting/"&gt;Cost Alerting &amp;amp; Budgets&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="the-mental-model"&gt;The mental model&lt;a class="anchor" href="#the-mental-model"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Google Cloud Storage (GCS) is &lt;strong&gt;object storage&lt;/strong&gt;. A &lt;strong&gt;bucket&lt;/strong&gt; is a globally named container; an &lt;strong&gt;object&lt;/strong&gt; is a file inside it. &lt;code&gt;gs://&lt;/code&gt; is GCS&amp;rsquo;s equivalent of a filesystem path:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/02-bigquery-ml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/02-bigquery-ml/</guid><description>&lt;h1 id="bigquery-ml"&gt;BigQuery ML&lt;a class="anchor" href="#bigquery-ml"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=6Kska20zQO4"&gt;&lt;img src="https://img.youtube.com/vi/6Kska20zQO4/0.jpg" alt="BigQuery ML: Machine Learning with Standard SQL" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;If your data is already in a warehouse, copying it into a notebook just to train a baseline is often the slowest and riskiest part of the job. BigQuery ML lets SQL create the first model where the table already lives.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~30 min hands-on
🔗 needs: SQL basics · &lt;a href="/2026-05/week-8/01-cloud-storage-ml/"&gt;Cloud Storage for ML&lt;/a&gt; · &lt;a href="/2026-05/week-7/09-cost-alerting/"&gt;Cost Alerting &amp;amp; Budgets&lt;/a&gt;&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/03-mlflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/03-mlflow/</guid><description>&lt;h1 id="mlflow"&gt;MLflow&lt;a class="anchor" href="#mlflow"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=kn51WgTTjCw"&gt;&lt;img src="https://img.youtube.com/vi/kn51WgTTjCw/0.jpg" alt="MLflow Python Tutorial — ML Model Experiment Tracking" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;“Best model” means nothing if nobody can answer: best compared with which run, trained on which data, with which parameters, and where is the model file? MLflow turns that scavenger hunt into an experiment record.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~25 min hands-on
🔗 needs: Python + &lt;code&gt;uv&lt;/code&gt; · &lt;a href="/2026-05/week-2/08-logging-testing/"&gt;Logging &amp;amp; Testing&lt;/a&gt; · &lt;a href="/2026-05/week-8/01-cloud-storage-ml/"&gt;Cloud Storage for ML&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="the-mental-model"&gt;The mental model&lt;a class="anchor" href="#the-mental-model"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;MLflow is an open-source platform for the ML lifecycle. Start with &lt;strong&gt;Tracking&lt;/strong&gt;: a local API and web UI that record what happened in each training attempt.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/04-finetuning-strategy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/04-finetuning-strategy/</guid><description>&lt;h1 id="fine-tuning-strategy"&gt;Fine-Tuning Strategy&lt;a class="anchor" href="#fine-tuning-strategy"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=L7PfLk4a2oY"&gt;&lt;img src="https://img.youtube.com/vi/L7PfLk4a2oY/0.jpg" alt="Fine-Tuning vs. RAG Explained in 4 Minutes" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Fine-tuning is not a button that makes an AI “know your business.” It is a costly way to change a model&amp;rsquo;s behaviour. Before you train, prove that prompting, retrieval, or tools cannot solve the real problem more safely and cheaply.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~30 min hands-on
🔗 needs: &lt;a href="/2026-05/week-3/01-prompt-engineering-1-foundations/"&gt;Prompt Engineering&lt;/a&gt; · &lt;a href="/2026-05/week-4/ragas-evaluation/"&gt;RAGAS Evaluation&lt;/a&gt; · &lt;a href="/2026-05/week-8/03-mlflow/"&gt;MLflow&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="what-fine-tuning-changes"&gt;What fine-tuning changes&lt;a class="anchor" href="#what-fine-tuning-changes"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A pretrained model has learned broad patterns from its original training. &lt;strong&gt;Fine-tuning&lt;/strong&gt; continues training it on examples for a narrower task. In supervised fine-tuning (SFT), each example demonstrates the input and the desired output.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/05-huggingface-ecosystem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/05-huggingface-ecosystem/</guid><description>&lt;h1 id="hugging-face-ecosystem"&gt;Hugging Face Ecosystem&lt;a class="anchor" href="#hugging-face-ecosystem"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=aeiUTRvh6yE"&gt;&lt;img src="https://img.youtube.com/vi/aeiUTRvh6yE/0.jpg" alt="Hands-On Hugging Face Tutorial" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Hugging Face is not “a website to download models.” It is an ecosystem for finding, running, evaluating, adapting, documenting, and sharing ML artefacts. The model page is part of the model. Read it before you run it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~30 min hands-on
🔗 needs: Python + &lt;code&gt;uv&lt;/code&gt; · &lt;a href="/2026-05/week-8/04-finetuning-strategy/"&gt;Fine-Tuning Strategy&lt;/a&gt; · &lt;a href="/2026-05/week-8/03-mlflow/"&gt;MLflow&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="the-map-of-the-ecosystem"&gt;The map of the ecosystem&lt;a class="anchor" href="#the-map-of-the-ecosystem"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Hub ── stores/version-controls ── models, datasets, Spaces
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; │
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── Transformers ── load models, tokenize, infer, train
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── Datasets ── load, clean, split, stream datasets
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── Evaluate ── calculate and share metrics
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── PEFT ── train small adapters such as LoRA
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── TRL ── supervised / preference / RL training helpers
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ├── Accelerate ── place training across available hardware
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; └── Gradio / Spaces ── make a shareable demo&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;You do not need every library for a first project. A beginner path is:&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/06-finetuning-techniques/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/06-finetuning-techniques/</guid><description>&lt;h1 id="fine-tuning-techniques"&gt;Fine-Tuning Techniques&lt;a class="anchor" href="#fine-tuning-techniques"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=L7PfLk4a2oY"&gt;&lt;img src="https://img.youtube.com/vi/L7PfLk4a2oY/0.jpg" alt="Fine-Tuning vs. RAG Explained" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;The important question is not “which fine-tuning buzzword should I use?” It is “what minimum change to the model, with what evidence, improves my held-out task?” Start with an adapter. Earn the right to do anything more expensive.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~14 min read · ~35 min hands-on
🔗 needs: &lt;a href="/2026-05/week-8/04-finetuning-strategy/"&gt;Fine-Tuning Strategy&lt;/a&gt; · &lt;a href="/2026-05/week-8/05-huggingface-ecosystem/"&gt;Hugging Face Ecosystem&lt;/a&gt; · &lt;a href="/2026-05/week-8/07-quantization/"&gt;Quantization&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="the-techniques-from-least-to-most-invasive"&gt;The techniques, from least to most invasive&lt;a class="anchor" href="#the-techniques-from-least-to-most-invasive"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Technique&lt;/th&gt;
					&lt;th&gt;What changes&lt;/th&gt;
					&lt;th&gt;Best first use&lt;/th&gt;
					&lt;th&gt;Main risk&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;Prompt / few-shot&lt;/td&gt;
					&lt;td&gt;no weights&lt;/td&gt;
					&lt;td&gt;establish a baseline&lt;/td&gt;
					&lt;td&gt;prompt becomes long/fragile&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;SFT&lt;/strong&gt; (supervised fine-tuning)&lt;/td&gt;
					&lt;td&gt;weights learn input → desired output examples&lt;/td&gt;
					&lt;td&gt;stable formatting, classification, extraction, task behaviour&lt;/td&gt;
					&lt;td&gt;copies mistakes in labels&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;LoRA&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;trains small low-rank adapter matrices; base model stays frozen&lt;/td&gt;
					&lt;td&gt;almost every first open-model adaptation&lt;/td&gt;
					&lt;td&gt;wrong base model/template still fails&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;QLoRA&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;LoRA while the frozen base model is loaded in 4-bit&lt;/td&gt;
					&lt;td&gt;GPU/RAM-constrained first run&lt;/td&gt;
					&lt;td&gt;hardware/software compatibility and quality trade-off&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Preference tuning (DPO etc.)&lt;/td&gt;
					&lt;td&gt;learns chosen output over rejected output&lt;/td&gt;
					&lt;td&gt;you can rank alternatives more easily than write a perfect one&lt;/td&gt;
					&lt;td&gt;unclear/inconsistent preferences&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Full fine-tuning&lt;/td&gt;
					&lt;td&gt;updates all/most base weights&lt;/td&gt;
					&lt;td&gt;specialised, well-funded work with strong data/evals&lt;/td&gt;
					&lt;td&gt;cost, catastrophic forgetting, hard rollback&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;Continued pretraining&lt;/td&gt;
					&lt;td&gt;learns from raw domain text before task tuning&lt;/td&gt;
					&lt;td&gt;a large, licensed domain corpus materially differs from base knowledge&lt;/td&gt;
					&lt;td&gt;costly; raw text is often a bad substitute for RAG&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For this course, the practical default is &lt;strong&gt;instruction-tuned base model + supervised fine-tuning + LoRA&lt;/strong&gt;, usually with 4-bit QLoRA. It is cheap enough to learn, produces a small adapter to share, and is reversible: remove the adapter to return to the base model.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/07-quantization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/07-quantization/</guid><description>&lt;h1 id="quantization"&gt;Quantization&lt;a class="anchor" href="#quantization"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=6TF_00jiivk"&gt;&lt;img src="https://img.youtube.com/vi/6TF_00jiivk/0.jpg" alt="QLoRA in Action: a Fine-Tuning Tutorial" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Quantization is a compression trade-off: use fewer bits to store model numbers so the model fits on available hardware. It makes a model cheaper to load; it does not make the model smaller in judgement, safer, or better at your task.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~25 min hands-on
🔗 needs: &lt;a href="/2026-05/week-8/06-finetuning-techniques/"&gt;Fine-Tuning Techniques&lt;/a&gt; · basic Python&lt;/p&gt;
&lt;h2 id="from-precision-to-practical-memory"&gt;From precision to practical memory&lt;a class="anchor" href="#from-precision-to-practical-memory"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Model weights are large arrays of numbers. Full precision uses many bits for each number; quantization represents them with fewer bits plus scales/metadata that approximately reconstruct the original values.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/08-gemma4-finetuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/08-gemma4-finetuning/</guid><description>&lt;h1 id="gemma-4-fine-tuning"&gt;Gemma 4 Fine-Tuning&lt;a class="anchor" href="#gemma-4-fine-tuning"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=QZSbRMIJlRw"&gt;&lt;img src="https://img.youtube.com/vi/QZSbRMIJlRw/0.jpg" alt="How to Fine-Tune Gemma 4 on Your Own Dataset with Unsloth" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Your first successful fine-tune should be small, boring, and measurable: a small instruction model, a narrow custom task, a LoRA adapter, and a held-out comparison. A giant model and a spectacular demo are not the same as evidence.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~15 min read · ~45 min hands-on
🔗 needs: &lt;a href="/2026-05/week-8/04-finetuning-strategy/"&gt;Fine-Tuning Strategy&lt;/a&gt; · &lt;a href="/2026-05/week-8/06-finetuning-techniques/"&gt;Fine-Tuning Techniques&lt;/a&gt; · &lt;a href="/2026-05/week-8/07-quantization/"&gt;Quantization&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="why-gemma-4--unsloth-is-a-good-first-lab"&gt;Why Gemma 4 + Unsloth is a good first lab&lt;a class="anchor" href="#why-gemma-4--unsloth-is-a-good-first-lab"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Gemma 4 is a family of Google open-weight models with instruction-tuned variants. &lt;strong&gt;Unsloth&lt;/strong&gt; provides guided notebooks and Studio workflows for loading compatible models, applying a chat template, performing LoRA/QLoRA fine-tuning, testing, and exporting an adapter.&lt;/p&gt;</description></item><item><title/><link>/2026-05/week-8/10-model-publishing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/week-8/10-model-publishing/</guid><description>&lt;h1 id="model-publishing--model-cards"&gt;Model Publishing &amp;amp; Model Cards&lt;a class="anchor" href="#model-publishing--model-cards"&gt;#&lt;/a&gt;&lt;/h1&gt;
&lt;blockquote class='book-hint '&gt;
&lt;p&gt;&lt;strong&gt;Uploading weights is not publishing a model. Publishing means another person can discover what it is, load the correct artefact, understand what it was tested for, avoid known harms, and decide whether they are allowed to use it. The model card is the interface for that decision.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;⏱ ~12 min read · ~30 min hands-on
🔗 needs: &lt;a href="/2026-05/week-8/05-huggingface-ecosystem/"&gt;Hugging Face Ecosystem&lt;/a&gt; · &lt;a href="/2026-05/week-8/08-gemma4-finetuning/"&gt;Gemma 4 Fine-Tuning&lt;/a&gt; · &lt;a href="/2026-05/week-8/03-mlflow/"&gt;MLflow&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="publish-an-artefact-not-a-surprise"&gt;Publish an artefact, not a surprise&lt;a class="anchor" href="#publish-an-artefact-not-a-surprise"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A release answers five questions before someone downloads it:&lt;/p&gt;</description></item><item><title/><link>/2026-05/wikipedia-data-with-python/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/2026-05/wikipedia-data-with-python/</guid><description>&lt;h2 id="wikipedia-data-with-python"&gt;Wikipedia Data with Python&lt;a class="anchor" href="#wikipedia-data-with-python"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://youtu.be/b6puvm-QEY0"&gt;&lt;img src="https://i.ytimg.com/vi_webp/b6puvm-QEY0/sddefault.webp" alt="Wikipedia data with Wikimedia Python library" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll learn how to scrape data from Wikipedia using the &lt;code&gt;wikipedia&lt;/code&gt; Python library, covering:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Installing and Importing&lt;/strong&gt;: Use pip install to get the Wikipedia library and import it with import wikipedia as wk.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keyword Search&lt;/strong&gt;: Use the search function to find Wikipedia pages containing a specific keyword, limiting results with the results argument.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fetching Summaries&lt;/strong&gt;: Use the summary function to get a concise summary of a Wikipedia page, limiting sentences with the sentences argument.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Retrieving Full Pages&lt;/strong&gt;: Use the page function to obtain the full content of a Wikipedia page, including sections and references.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accessing URLs&lt;/strong&gt;: Retrieve the URL of a Wikipedia page using the url attribute of the page object.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracting References&lt;/strong&gt;: Use the references attribute to get all reference links from a Wikipedia page.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fetching Images&lt;/strong&gt;: Access all images on a Wikipedia page via the images attribute, which returns a list of image URLs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracting Tables&lt;/strong&gt;: Use the pandas.read_html function to extract tables from the HTML content of a Wikipedia page, being mindful of table indices.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are links and references:&lt;/p&gt;</description></item></channel></rss>