正在加载内容...

963963 Chat Guide Portal Independent coverage of news

Site Topics Fundamentals 5 in Practice: Lessons From Real Deployments

By Laura Bennett · · 1146 words
Site Topics Fundamentals 5 in Practice: Lessons From Real Deployments

Cloud Infrastructure: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Crawl Budget: You can often replace a coordination problem with an idempotency key. Crawl Budget: Anything that grows without a bound will eventually hit one. Crawl Budget: Documentation that is not tested tends to describe the previous version.

Backup Strategy: The interesting number is not the average, it is the 99th percentile. Backup Strategy: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Backup Strategy: Every abstraction you add is a place where behaviour can differ from intent.

Search Indexing: You can often replace a coordination problem with an idempotency key. Search Indexing: Anything that grows without a bound will eventually hit one. Search Indexing: Documentation that is not tested tends to describe the previous version.

Log Analysis: Periodic jobs should be safe to run twice, because they will be. Log Analysis: You rarely need a new component to fix a boundary problem. Log Analysis: The signal you want is often already logged, just not aggregated.

Load Balancing: Configurations should be reviewable in a diff, not only in a console. Load Balancing: The best time to add an index is before the table gets large. Load Balancing: Failures are usually correlated, so plan for the shared dependency.

Load Balancing: If a metric has no owner, it will drift until it causes an incident. Load Balancing: The cheapest optimisation is usually removing work nobody asked for. Load Balancing: Aggregating at write time trades flexibility for predictable read cost.

Configurations should be reviewable in a diff, not only in a console. This is most visible in load balancing. Consider load balancing specifically. The best time to add an index is before the table gets large. Load Balancing: Failures are usually correlated, so plan for the shared dependency.

Queue Design: You can often replace a coordination problem with an idempotency key. Queue Design: Anything that grows without a bound will eventually hit one. Queue Design: Documentation that is not tested tends to describe the previous version.

API Design: If the rollback plan needs a meeting, it is not a rollback plan. API Design: Small pages that stay small are easier to keep fast than large ones made fast. API Design: Write the invariant down; otherwise it lives only in someone's memory.

Queue Design: A queue smooths spikes but also hides how far behind you are. Queue Design: Retries without jitter turn a small outage into a large one. Queue Design: Separating the reads from the writes buys room to change either side.

For search indexing, the constraint matters more than the feature list. Configurations should be reviewable in a diff, not only in a console. Teams working on search indexing usually discover this the hard way. The best time to add an index is before the table gets large. Failures are usually correlated, so plan for the shared dependency. This is most visible in search indexing.

A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for cloud infrastructure. For cloud infrastructure, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on cloud infrastructure usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.

Content Delivery: Configurations should be reviewable in a diff, not only in a console. Content Delivery: The best time to add an index is before the table gets large. Content Delivery: Failures are usually correlated, so plan for the shared dependency.

Content Delivery: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to content delivery as well. In practice, content delivery behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Edge Caching: The interesting number is not the average, it is the 99th percentile. Edge Caching: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Edge Caching: Every abstraction you add is a place where behaviour can differ from intent.

A queue smooths spikes but also hides how far behind you are. This is most visible in search indexing. Consider search indexing specifically. Retries without jitter turn a small outage into a large one. Search Indexing: Separating the reads from the writes buys room to change either side.

Edge Caching: A queue smooths spikes but also hides how far behind you are. Edge Caching: Retries without jitter turn a small outage into a large one. Edge Caching: Separating the reads from the writes buys room to change either side.

Storage Tiers: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to storage tiers as well. In practice, storage tiers behaves differently: Failures are usually correlated, so plan for the shared dependency.

Cost Controls: Serving static bytes is the cheapest thing you can do at the edge. Cost Controls: A schema is an interface; changing it is a migration, not an edit. Cost Controls: Track the denominator as carefully as the numerator.

Teams working on content delivery usually discover this the hard way. If the rollback plan needs a meeting, it is not a rollback plan. Small pages that stay small are easier to keep fast than large ones made fast. This is most visible in content delivery. Consider content delivery specifically. Write the invariant down; otherwise it lives only in someone's memory.

Observability: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to observability as well. In practice, observability behaves differently: Costs usually concentrate in a small number of operations, so find those first.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for access control. For access control, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on access control usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Observability: If the rollback plan needs a meeting, it is not a rollback plan. Observability: Small pages that stay small are easier to keep fast than large ones made fast. Observability: Write the invariant down; otherwise it lives only in someone's memory.

Related reading