1. Keep services running: Runit, even with daemontools differences, is hard to beat.
2. Hard(-ish) resource limits and accounting: LXC w/ cgroups. Almost as good as full paravirtualization (Xen). There are still some issues with limiting resource contention impact between cgroups.
3. Softer resource limits, rogue app restarter: We've heavily modified bluepill because it seemed to lack insight on the needs and challenges of large-scale production ops. Specifically, we've added optional total child process limits (a few issues reported, fixed and even submitted a pull request). It might be useful to add max # of processes, nic bandwidth, iops and couple other checks.
2. Hard(-ish) resource limits and accounting: LXC w/ cgroups. Almost as good as full paravirtualization (Xen). There are still some issues with limiting resource contention impact between cgroups.
3. Softer resource limits, rogue app restarter: We've heavily modified bluepill because it seemed to lack insight on the needs and challenges of large-scale production ops. Specifically, we've added optional total child process limits (a few issues reported, fixed and even submitted a pull request). It might be useful to add max # of processes, nic bandwidth, iops and couple other checks.
Also worth considering:
4. Status monitoring: Icinga
5. Performance: collectd
6. Entropy injection: Chaos Monkey