Software support · Incident response · Enterprise systems

Practical lessons from complex production environments.

Field notes on technical operations, reliability and the decisions that matter during incidents and infrastructure changes.

Articles

Incident response · Distributed systems · Resilience engineering

How Retry Storms Can Turn Recovery into a Self-Inflicted Denial-of-Service Incident

Why synchronised retries, overloaded monitoring and cascading failures can prevent a distributed platform from recovering.

Read article →
SQL Server · Windows Failover Cluster · High availability

When a Network Incident Silently Affects Your SQL Server Cluster

Why restoring connectivity is not enough when cluster membership, replicas and reporting services may remain degraded.

Read article →
SQL Server · Change management · Upgrade strategy

Why In-Place SQL Server Upgrades Rarely Behave the Same Twice

The runbook may be identical, but the actual state and installation history of each server usually are not.

Read article →

About

About the publication

Ops Field Notes is an independent technical publication focused on enterprise systems, SQL Server, incident response, reliability and operational decision-making. The articles are intentionally abstracted to protect confidential information and avoid identifying specific organisations or production environments.