Sunday, February 1, 2009

Data Center Failure As A Learning Experience

This short article emphasizes the need to plan for failure when designing a company's data center (or data centers). The main idea, which is somewhat provocative at first sight, is that you can't design for 100% uptime so you might as well ensure that failure is "small" (isolated) and recoverable. This is somewhat reminiscent of Eric Brewer's point in his "lessons from large-scale systems" work that mean time to recover (MTTR) is more important than mean time to failure (MTTF).

The other interesting idea in the article is that hardware failures tend to happen either shortly after installing hardware ("infant mortality") or as it wears down towards the end of its life cycle ("old age"). Thus it's important to wear in newly bought hardware for some amount of time to weed out the equipment that will fail early.

No comments: