AI costs · Evals · Production engineering

    Notes from the deployment layer.

    Independent FDE perspectives, model updates, and practical guides to costs, evaluations, and production AI.

    All articles

    Newest first

    01

    Top 10 AI Transformation Companies to Consider in 2026

    A researched editorial shortlist of ten AI transformation providers, with Deployed Engineer featured first as the publisher's practice, service comparisons, and questions to ask before choosing a partner.

    AI Transformation8 min
    02

    Top 10 AI Transformation Companies: A 2025 Retrospective and Industry Playbook

    A 2025-informed editorial shortlist, published in 2026, featuring Deployed Engineer and nine providers, with six clearly hypothetical industry case studies.

    AI Transformation8 min
    03

    Jev for Classification: Six Places to Use It in Real Workflows

    A researched guide to TypeSafe's Jev with six classification workflows, a Python ticket-routing example, and a practical evaluation plan.

    Jev8 min
    04

    Independent FDE: Bringing AI Into the Work Your Business Runs On

    The opportunity in AI services sits between model capability and everyday operations. Here is how an independent forward deployed engineer closes that gap.

    Independent FDE4 min
    05

    September Model Updates: OpenAI, Grok, Gemini, and Claude Through an FDE Lens

    GPT-6.1 Sol, Grok 4.7, Gemini 3.8 Flash, and Claude Opus 5.5 bring new options for production workflows. A source-linked briefing on what changed and what to evaluate.

    Model Updates4 min
    06

    How to Reduce AI Costs Without Breaking Quality

    A first-principles playbook for cutting LLM and AI infrastructure costs by measuring cost per successful task, routing work, caching context, and using evals as the safety rail.

    Reduce AI Costs5 min
    07

    How to Evaluate a Simple AI Chatbot

    A beginner-friendly guide to testing a support chatbot with five questions, simple pass-or-fail checks, and a repeatable spreadsheet workflow.

    Chatbot Evals6 min
    08

    How to Test a Chatbot That Calls Tools

    A simple guide to checking whether a chatbot chose the right tool, used the right details, asked for confirmation, and handled failures safely.

    Tool Calling Evals4 min
    09

    How to Evaluate a Simple RAG Chatbot

    Test whether a RAG chatbot found the right passage, answered from that passage, cited it, and admitted when the answer was missing.

    RAG Evals5 min
    10

    GPT-6 Astra: Where It Helps, What It Costs, and How to Use It Well

    A practical, source-linked guide to GPT-6 Astra's long-context reasoning, computer use, coding, and professional workflows—with a cost-aware way to decide when it is actually the right model.

    GPT-6 Astra5 min
    11

    Build a Useful Agent with OpenAI’s Agents API: A Small Project That Teaches the Big Ideas

    A step-by-step demo for building a support-triage agent with tools, a sandbox, approvals, subagents, and traces—without hiding the important engineering decisions behind a framework.

    Agents API6 min
    12

    AI Agent News, August 2026: MCP, Grok Bot, DeepSeek Harness, and Cursor Origin

    A source-checked briefing on the launches moving agents from chat windows into browsers, developer infrastructure, and long-running company workflows.

    AI Agents3 min
    13

    Reading Code Is Getting Harder. Learn the System Instead

    When agents can generate more code than a person can read, understanding has to come from traces, executable examples, diagrams, and sequential explanations.

    Software Engineering3 min
    14

    AI Search Visibility: What Gamma's AEO Playbook Gets Right

    A practical reading of Gamma's reported AI visibility tactics, including first-party content, agent-readable actions, comparison pages, and citation monitoring.

    AI Search3 min
    15

    Cursor Origin, Command Code, and the New Coding Agent Stack

    Cursor is expanding from editor to forge while Command Code learns from developer choices. Here is what changed, including the verified Grok model and current pricing.

    Cursor3 min
    16

    Qwen 3.7, GLM-5.3, and the Gemini 3.7 Rumor: A Verified Model Update

    A source-checked model briefing on Qwen3.7 Flash, Qwen-Image 3.0, GLM-5.3, and the model name that Google has not announced.

    Qwen2 min
    17

    Spotify Launched an AI Agent. The Interesting Part Is Taste

    Studio by Spotify Labs can research, organize work, use connected context, and turn it into personal audio. That makes taste part of the agent interface.

    Spotify2 min
    18

    Unsloth Desktop Makes Local AI Feel Like a Product

    Unsloth brings local model use, comparison, tool calling, file analysis, and training into an offline desktop workflow for Mac and Windows.

    Unsloth2 min
    19

    Claude Code Session Costs and OpenAI's GPT-5.6 Price Drop

    Anthropic explains how session habits affect token use while OpenAI cuts GPT-5.6 Luna and Terra prices. Both point toward cost per successful task.

    Claude Code3 min
    20

    YC's Org Code Idea: Designing Companies for Agents

    A YC launch for Pentagon puts a useful name on an emerging practice: agent teams need explicit, testable organization design.

    Y Combinator3 min
    21

    AI Engineer World's Fair 2026 Notes: Computer Use, Evals, GTM Agents, and Context

    My organized notes from AI Engineer World's Fair sessions on July 1 and July 2, covering computer use agents, eval design, agent simulations, CaaS, AI-first SDLC, LLM wikis, and Ramp's GTM agent platform.

    AI Engineering15 min
    22

    How Ramp Does Forward Deployed Engineering

    Notes from Leo Mehr on Ramp's FDE practice: always be scoping, interrogate urgency, validate assumptions, and use AI to speed up request triage.

    Forward Deployed Engineering2 min
    23

    Anthropic's Forward Deployed Engineering 101

    Notes from Kevin Bai on Anthropic's FDE framing: sell outcomes, build on shared primitives, and only add FDE when a complex product meets a non-technical buyer.

    Forward Deployed Engineering2 min
    24

    Claude Fable 5 Is Back: What Anthropic's Mythos-Class Launch Means for Agent Engineers

    Anthropic launched Claude Fable 5 on June 9, suspended access on June 12 after US export controls, then restored global access on July 1. Here is what changed, what the safeguards mean, and where Fable fits in serious agent stacks.

    LLM Engineering4 min
    25

    GLM-5.2 Just Made Open-Weight Agent Models Harder to Ignore

    Z.ai released GLM-5.2 with a solid 1M-token context window, MIT-licensed weights, 744B total parameters with 40B active, and strong long-horizon coding results. The open-weight agent stack is getting serious.

    Open Source4 min
    26

    Vercel's eve Turns the Agent Harness Into a Directory

    Vercel launched eve, an open-source framework for production agents. Instructions live in Markdown, tools in TypeScript, skills and subagents in folders, with durable execution, sandboxed compute, approvals, and evals built in.

    Agent Engineering3 min
    27

    Loop Engineering Took Over X. The Useful Part Is the Stop Condition

    Loop engineering moved from X posts into serious engineering blogs in June. LangChain's stacked-loop framing makes the idea more concrete: agent loops, verification loops, event-driven loops, and trace-driven improvement loops.

    Agent Engineering11 min
    28

    Claude Opus 4.7 Is Here: The Benchmarks, the Pricing, and What Actually Matters

    Anthropic shipped Claude Opus 4.7 this week. 87.6% on SWE-bench Verified, 64.3% on SWE-bench Pro, a 1M-token context window, and the same $5/$25 pricing as Opus 4.6. Here is my honest take on what changed, what did not, and whether it deserves a spot in your stack.

    LLM Engineering6 min
    29

    OpenAI's Codex Scratchpad: Parallel Agents Are Becoming Table Stakes

    OpenAI is wiring a Scratchpad UI into the Codex app that fires multiple agents in parallel from a TODO list. A small feature with a big signal. Parallel agent execution is no longer a differentiator. It is the default.

    Agent Engineering3 min
    30

    Google's Gemma 4 Shipped Apache 2.0: Why Open Weights Just Got Serious

    Google released Gemma 4 on April 2 under Apache 2.0. Four sizes from 2B to 31B, 256K context, native vision and audio, and real agentic capabilities. The open-model landscape just changed, and not everyone is ready for what that means.

    LLM Engineering3 min
    31

    What is OpenClaw? The Open-Source AI Agent Taking Over GitHub

    OpenClaw is an open-source autonomous AI agent with 310K+ GitHub stars. Learn how it bridges messaging apps and AI models to automate tasks locally — and why agent engineers should care.

    Agent Engineering12 min
    32

    GPT-4.1, o3, and o4-mini: A Guide to OpenAI's Latest Models

    OpenAI shipped GPT-4.1, o3, o3-pro, and o4-mini. Here is what each model does, when to use it, and how they change the AI engineering stack.

    LLM Engineering12 min
    33

    Claude Code Tips and Tricks for AI Engineers (2026 Edition)

    The 2026 guide to Claude Code: voice mode, /loop, skills system, 1M context, Opus 4.6, and the workflows that top AI engineers actually use.

    AI Engineering11 min
    34

    Cursor AI in 2026: Automations, Parallel Agents, and MCP Apps

    Cursor 2026 adds event-driven automations, parallel cloud agents, Bugbot Autofix, and MCP Apps. Here is what changed and why it matters for AI engineers.

    AI Engineering10 min
    35

    MCP Servers Directory: Find the Right Model Context Protocol Integration

    A complete guide to MCP directories with 1200+ servers for AI engineers building with Claude Code and Cursor. The Model Context Protocol ecosystem mapped.

    Context Engineering10 min
    36

    AI CLI Tools: Claude Code vs. Codex vs. Gemini CLI vs. Aider vs. OpenCode

    A direct comparison of the 5 best AI CLI tools for developers in 2026 — pricing, capabilities, and which to choose for your engineering workflow.

    AI Engineering12 min
    37

    What is a Forward Deployed Engineer? The Complete Guide

    A forward deployed engineer (FDE) embeds directly with customers to solve their hardest technical problems. Learn about the role's origin at Palantir, why it grew 800% in 2025, salary data ($238K-$630K+), required skills, and how to break into this career as a forward deployed AI engineer.

    Forward Deployed Engineering11 min
    38

    Cursor AI Tips That Will 10x Your Coding Speed in 2025

    Master Cursor AI with these viral tips that are transforming how engineers write code. From prompt engineering to code generation, learn the secrets that top developers are using.

    AI Engineering4 min
    39

    Claude Coding Tips: The Secret Techniques Top Engineers Are Using

    Discover the viral Claude coding techniques that are revolutionizing how developers approach problem-solving and code generation.

    AI Engineering4 min
    40

    The Ultimate Engineering Learning Resources Guide (2025)

    Curated list of the best learning resources for engineers in 2025. From AI tools to system design, discover what top engineers are using to stay ahead.

    AI Engineering5 min