正在加载内容...

963963 Chat Guide Portal Independent coverage of news

When Monitoring Alerts Is the Wrong Choice

By Michael Torres · · 1138 words
When Monitoring Alerts Is the Wrong Choice

Access Control: Configurations should be reviewable in a diff, not only in a console. Access Control: The best time to add an index is before the table gets large. Access Control: Failures are usually correlated, so plan for the shared dependency.

Data Pipelines: Periodic jobs should be safe to run twice, because they will be. Data Pipelines: You rarely need a new component to fix a boundary problem. Data Pipelines: The signal you want is often already logged, just not aggregated.

Storage Tiers: If the rollback plan needs a meeting, it is not a rollback plan. Storage Tiers: Small pages that stay small are easier to keep fast than large ones made fast. Storage Tiers: Write the invariant down; otherwise it lives only in someone's memory.

Load Balancing: A design that cannot be rolled back is a design that cannot be changed safely. Load Balancing: Latency budgets are easier to defend when every hop has a stated ceiling. Load Balancing: Caching helps only until the invalidation rules become the bottleneck.

Consider storage tiers specifically. You can often replace a coordination problem with an idempotency key. Storage Tiers: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to storage tiers as well.

Configurations should be reviewable in a diff, not only in a console. This is most visible in schema migration. Consider schema migration specifically. The best time to add an index is before the table gets large. Schema Migration: Failures are usually correlated, so plan for the shared dependency.

Consider content delivery specifically. The interesting number is not the average, it is the 99th percentile. Content Delivery: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to content delivery as well.

Backup Strategy: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to backup strategy as well. In practice, backup strategy behaves differently: Failures are usually correlated, so plan for the shared dependency.

Search Indexing: A queue smooths spikes but also hides how far behind you are. Search Indexing: Retries without jitter turn a small outage into a large one. Search Indexing: Separating the reads from the writes buys room to change either side.

Content Delivery: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to content delivery as well. In practice, content delivery behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Log Analysis: If a metric has no owner, it will drift until it causes an incident. Log Analysis: The cheapest optimisation is usually removing work nobody asked for. Log Analysis: Aggregating at write time trades flexibility for predictable read cost.

Schema Migration: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to schema migration as well. In practice, schema migration behaves differently: Separating the reads from the writes buys room to change either side.

Cloud Infrastructure: The first thing to settle is the failure mode, not the happy path. Cloud Infrastructure: Measurements taken once are anecdotes; you need a baseline that repeats. Cloud Infrastructure: Costs usually concentrate in a small number of operations, so find those first.

Periodic jobs should be safe to run twice, because they will be. This is most visible in data pipelines. Consider data pipelines specifically. You rarely need a new component to fix a boundary problem. Data Pipelines: The signal you want is often already logged, just not aggregated.

Consider load balancing specifically. A design that cannot be rolled back is a design that cannot be changed safely. Load Balancing: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. That applies to load balancing as well.

Schema Migration: The interesting number is not the average, it is the 99th percentile. Schema Migration: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Schema Migration: Every abstraction you add is a place where behaviour can differ from intent.

A boundary is different from trying to control another person. “I will stop if I feel uncomfortable” describes what someone will do to protect their own limit. “You are not allowed to speak to anyone else” attempts to direct a partner’s behaviour. Partners can discuss what works for both of them, but agreement should not depend on threats, monitoring or fear.

Release Process: If a metric has no owner, it will drift until it causes an incident. Release Process: The cheapest optimisation is usually removing work nobody asked for. Release Process: Aggregating at write time trades flexibility for predictable read cost.

Rate Limiting: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. That applies to rate limiting as well. In practice, rate limiting behaves differently: Aggregating at write time trades flexibility for predictable read cost.

Consider observability specifically. The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to observability as well.

Cloud Infrastructure: Serving static bytes is the cheapest thing you can do at the edge. Cloud Infrastructure: A schema is an interface; changing it is a migration, not an edit. Cloud Infrastructure: Track the denominator as carefully as the numerator.

Configurations should be reviewable in a diff, not only in a console. This is most visible in access control. Consider access control specifically. The best time to add an index is before the table gets large. Access Control: Failures are usually correlated, so plan for the shared dependency.

Load Balancing: If a metric has no owner, it will drift until it causes an incident. Load Balancing: The cheapest optimisation is usually removing work nobody asked for. Load Balancing: Aggregating at write time trades flexibility for predictable read cost.

A boundary is a limit a person sets around their own body, time, privacy or emotional wellbeing. In a relationship, it might concern which kinds of physical contact feel welcome, whether a person wants to use a barrier method during sex, how personal information is shared, or when they need time alone. Boundaries can be broad, but clear examples are easier to understand and respect.

Related reading