Field notes
What production behavior teaches us.
Practical guides to MCP distribution, analytics, instrumentation, and agent observability.
pick the question you are working on
MCP Observability: How to Monitor an MCP Server in Production
Learn how to monitor MCP servers across infrastructure, tool calls, sessions, user behavior, outcomes, and recurring production failures.
Read the field note ↗MCP Analytics: How to See What Users Are Actually Doing With Your MCP Server
Measure MCP users, use cases, journeys, outcomes, unmet needs, feedback, and commercial signals instead of relying on tool-call volume alone.
Buyer's guide · 12 minBest MCP Observability and Monitoring Tools in 2026
Compare Flowlines, Grafana, Sentry, Datadog, OpenLIT, TrackMCP, and Honeycomb by infrastructure, traces, tool analytics, sessions, behavior, and outcomes.
MCP strategy · 10 minWhen Should You Build an MCP Server? A Practical Decision Framework
Decide whether your product needs an MCP server, a direct API integration, an agent skill, static context, or browser automation, then validate the business case.
MCP engineering · 10 minMCP OpenTelemetry Instrumentation: How to Instrument an MCP Server
Learn how to instrument an MCP server with OpenTelemetry, trace tool calls, propagate context, protect sensitive data, and validate telemetry.
MCP distribution · 15 minHow to Get Listed in the OpenAI Plugins Directory: MCP Submission Guide
Let Codex prepare your OpenAI plugin submission with Plugin Creator, then complete the MCP tests, listing, policies, and review.
MCP distribution · 11 minHow to Get Listed in the Claude Plugin Marketplace: Complete Submission Guide
Build, validate, submit, and distribute a Claude plugin, with the current public-repository requirements, package structure, review path, and MCP launch checklist.
Guide · 5 minHow to monitor AI agents in production
A practical framework for monitoring production AI agents across execution, quality, behavior, users, and outcomes, with a path back to original sessions.
Evaluation · 4 minHow to evaluate AI agents in production
Evaluate production agents with outcome metrics, behavioral signals, cohort analysis and sampled human review without confusing an LLM judge with ground truth.
Engineering · 4 minHow to debug AI agents in production
Debug production agents by moving from an outcome gap to a recurring behavior, affected cohort, exact trace and verified release change.
Guide · 6 minAI Agent Failure Modes in Production: The Complete Taxonomy (2026)
A working taxonomy of production AI agent failure modes: false success, loops, silent drift, intent mismatch, and cohort gaps, with what each looks like in a real session.
Guide · 10 minThe 9 Best AI Agent Observability Tools in 2026
Compare nine AI agent observability tools across tracing, evaluation, production sessions, cross-session analysis, OpenTelemetry, and MCP workflows.
Architecture · 2 minWhat is behavioral observability?
Behavioral observability explains how an AI agent behaves across production sessions, users, intents, issues, tools, and outcomes.
Engineering · 2 minHow to detect AI agent drift in production
Detect agent drift as recurring behavior across sessions, users, intents, and MCP journeys, then inspect the affected production sessions.
Engineering · 3 minHow to connect Flowlines to production traces
Connect OpenTelemetry, Langfuse, or LangSmith to Flowlines, verify the first session, and analyze behavior across production traces.
Opinion · 2 minFlowlines and Mem0 solve different production problems
Mem0 provides agent memory. Flowlines provides behavioral observability across production sessions, users, issues, and MCP journeys.
Product · 2 minIntent engineering requires production evidence
Intent taxonomies become useful when they are grounded in complete production sessions, connected to outcomes, and reviewed as user needs evolve.
Architecture · 2 minLLMs are stateless. Production agent systems should preserve context.
A production agent needs explicit session, user, release and tool context so teams can reconstruct behavior and evaluate outcomes over time.
Reliability · 2 minThe silent failure problem in AI agents
Silent agent failures look technically successful while the user outcome is wrong, incomplete or abandoned. Here is how to detect them in production.
Product · 2 minWhy AI agents do not improve from production data by themselves
Production agents do not learn automatically. Improvement requires observable outcomes, reviewed failure patterns, controlled changes and release verification.