Backup Strategy in Practice: Lessons From Real Deployments
Storage Tiers: You can often replace a coordination problem with an idempotency key. Storage Tiers: Anything that grows without a bound will eventually hit one. Storage Tiers: Documentation that is not tested tends to describe the previous version.
Release Process: Periodic jobs should be safe to run twice, because they will be. Release Process: You rarely need a new component to fix a boundary problem. Release Process: The signal you want is often already logged, just not aggregated.
Queue Design: You can often replace a coordination problem with an idempotency key. Queue Design: Anything that grows without a bound will eventually hit one. Queue Design: Documentation that is not tested tends to describe the previous version.
Consider content delivery specifically. The interesting number is not the average, it is the 99th percentile. Content Delivery: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to content delivery as well.
Teams working on backup strategy usually discover this the hard way. The interesting number is not the average, it is the 99th percentile. Adding a cache in front of a slow query is a fix; fixing the query is a cure. This is most visible in backup strategy. Consider backup strategy specifically. Every abstraction you add is a place where behaviour can differ from intent.
Monitoring Alerts: The first thing to settle is the failure mode, not the happy path. Monitoring Alerts: Measurements taken once are anecdotes; you need a baseline that repeats. Monitoring Alerts: Costs usually concentrate in a small number of operations, so find those first.
Cloud Infrastructure: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: The signal you want is often already logged, just not aggregated.
Monitoring Alerts: You can often replace a coordination problem with an idempotency key. Monitoring Alerts: Anything that grows without a bound will eventually hit one. Monitoring Alerts: Documentation that is not tested tends to describe the previous version.
Consider schema migration specifically. Serving static bytes is the cheapest thing you can do at the edge. Schema Migration: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. That applies to schema migration as well.
Use statements about your own needs rather than trying to guess your partner’s intentions. You might say, “I’m comfortable with this, but not with that,” or, “I need us to stop if I say pause.” Be specific about what you mean by words such as “slow down” or “check in.” Ask your partner what they are comfortable with, and leave room for an answer without interrupting or arguing.
You can also state your own boundaries. Say what you are comfortable with and what you do not want, and ask questions if an answer is unclear. Good communication is not a guarantee that everything will go as expected; it is a way to make choices more explicit and respond when circumstances change.
Schema Migration: You can often replace a coordination problem with an idempotency key. Schema Migration: Anything that grows without a bound will eventually hit one. Schema Migration: Documentation that is not tested tends to describe the previous version.
Many screens can be completed with urine, blood or self-collected swabs. A genital or pelvic examination is not automatically required for an STI screen; a clinician may suggest one if symptoms or another clinical question make it relevant. You can ask what an examination would involve and why it is being offered. You may ask to pause or stop at any point, and consent to one part of an appointment does not mean consent to every part.
Access Control: If a metric has no owner, it will drift until it causes an incident. Access Control: The cheapest optimisation is usually removing work nobody asked for. Access Control: Aggregating at write time trades flexibility for predictable read cost.
Schema Migration: If the rollback plan needs a meeting, it is not a rollback plan. Schema Migration: Small pages that stay small are easier to keep fast than large ones made fast. Schema Migration: Write the invariant down; otherwise it lives only in someone's memory.
Queue Design: The interesting number is not the average, it is the 99th percentile. Queue Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Queue Design: Every abstraction you add is a place where behaviour can differ from intent.
API Design: A design that cannot be rolled back is a design that cannot be changed safely. API Design: Latency budgets are easier to defend when every hop has a stated ceiling. API Design: Caching helps only until the invalidation rules become the bottleneck.
Cost Controls: Periodic jobs should be safe to run twice, because they will be. Cost Controls: You rarely need a new component to fix a boundary problem. Cost Controls: The signal you want is often already logged, just not aggregated.
Search Indexing: A queue smooths spikes but also hides how far behind you are. Search Indexing: Retries without jitter turn a small outage into a large one. Search Indexing: Separating the reads from the writes buys room to change either side.
Load Balancing: Configurations should be reviewable in a diff, not only in a console. Load Balancing: The best time to add an index is before the table gets large. Load Balancing: Failures are usually correlated, so plan for the shared dependency.
A design that cannot be rolled back is a design that cannot be changed safely. The same reasoning holds for queue design. For queue design, the constraint matters more than the feature list. Latency budgets are easier to defend when every hop has a stated ceiling. Teams working on queue design usually discover this the hard way. Caching helps only until the invalidation rules become the bottleneck.
Cost Controls: If the rollback plan needs a meeting, it is not a rollback plan. Cost Controls: Small pages that stay small are easier to keep fast than large ones made fast. Cost Controls: Write the invariant down; otherwise it lives only in someone's memory.
You can often replace a coordination problem with an idempotency key. That applies to queue design as well. In practice, queue design behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for queue design.
If a metric has no owner, it will drift until it causes an incident. This is most visible in content delivery. Consider content delivery specifically. The cheapest optimisation is usually removing work nobody asked for. Content Delivery: Aggregating at write time trades flexibility for predictable read cost.