Field notes

What production behavior teaches us.

Practical guides to MCP distribution, analytics, instrumentation, and agent observability.

pick the question you are working on

MCP observability · 12 min

MCP Observability: How to Monitor an MCP Server in Production

Learn how to monitor MCP servers across infrastructure, tool calls, sessions, user behavior, outcomes, and recurring production failures.

Read the field note ↗
20 field notes
MCP analytics · 10 min

MCP Analytics: How to See What Users Are Actually Doing With Your MCP Server

Measure MCP users, use cases, journeys, outcomes, unmet needs, feedback, and commercial signals instead of relying on tool-call volume alone.

Sep 12, 2026›
Buyer's guide · 12 min

Best MCP Observability and Monitoring Tools in 2026

Compare Flowlines, Grafana, Sentry, Datadog, OpenLIT, TrackMCP, and Honeycomb by infrastructure, traces, tool analytics, sessions, behavior, and outcomes.

Sep 25, 2026›
MCP strategy · 10 min

When Should You Build an MCP Server? A Practical Decision Framework

Decide whether your product needs an MCP server, a direct API integration, an agent skill, static context, or browser automation, then validate the business case.

Sep 12, 2026›
MCP engineering · 10 min

MCP OpenTelemetry Instrumentation: How to Instrument an MCP Server

Learn how to instrument an MCP server with OpenTelemetry, trace tool calls, propagate context, protect sensitive data, and validate telemetry.

Sep 25, 2026›
MCP distribution · 15 min

How to Get Listed in the OpenAI Plugins Directory: MCP Submission Guide

Let Codex prepare your OpenAI plugin submission with Plugin Creator, then complete the MCP tests, listing, policies, and review.

Sep 12, 2026›
MCP distribution · 11 min

How to Get Listed in the Claude Plugin Marketplace: Complete Submission Guide

Build, validate, submit, and distribute a Claude plugin, with the current public-repository requirements, package structure, review path, and MCP launch checklist.

Sep 12, 2026›
Guide · 5 min

How to monitor AI agents in production

A practical framework for monitoring production AI agents across execution, quality, behavior, users, and outcomes, with a path back to original sessions.

Sep 12, 2026›
Evaluation · 4 min

How to evaluate AI agents in production

Evaluate production agents with outcome metrics, behavioral signals, cohort analysis and sampled human review without confusing an LLM judge with ground truth.

Sep 12, 2026›
Engineering · 4 min

How to debug AI agents in production

Debug production agents by moving from an outcome gap to a recurring behavior, affected cohort, exact trace and verified release change.

Sep 12, 2026›
Guide · 6 min

AI Agent Failure Modes in Production: The Complete Taxonomy (2026)

A working taxonomy of production AI agent failure modes: false success, loops, silent drift, intent mismatch, and cohort gaps, with what each looks like in a real session.

Sep 4, 2026›
Guide · 10 min

The 9 Best AI Agent Observability Tools in 2026

Compare nine AI agent observability tools across tracing, evaluation, production sessions, cross-session analysis, OpenTelemetry, and MCP workflows.

Sep 4, 2026›
Architecture · 2 min

What is behavioral observability?

Behavioral observability explains how an AI agent behaves across production sessions, users, intents, issues, tools, and outcomes.

Sep 1, 2026›
Engineering · 2 min

How to detect AI agent drift in production

Detect agent drift as recurring behavior across sessions, users, intents, and MCP journeys, then inspect the affected production sessions.

Sep 1, 2026›
Engineering · 3 min

How to connect Flowlines to production traces

Connect OpenTelemetry, Langfuse, or LangSmith to Flowlines, verify the first session, and analyze behavior across production traces.

Sep 25, 2026›
Opinion · 2 min

Flowlines and Mem0 solve different production problems

Mem0 provides agent memory. Flowlines provides behavioral observability across production sessions, users, issues, and MCP journeys.

Sep 1, 2026›
Product · 2 min

Intent engineering requires production evidence

Intent taxonomies become useful when they are grounded in complete production sessions, connected to outcomes, and reviewed as user needs evolve.

Sep 4, 2026›
Architecture · 2 min

LLMs are stateless. Production agent systems should preserve context.

A production agent needs explicit session, user, release and tool context so teams can reconstruct behavior and evaluate outcomes over time.

Sep 4, 2026›
Reliability · 2 min

The silent failure problem in AI agents

Silent agent failures look technically successful while the user outcome is wrong, incomplete or abandoned. Here is how to detect them in production.

Sep 4, 2026›
Product · 2 min

Why AI agents do not improve from production data by themselves

Production agents do not learn automatically. Improvement requires observable outcomes, reviewed failure patterns, controlled changes and release verification.

Sep 4, 2026›