I don't currently have anything hosted at Github, but I use them regularly to check out other people's code and their uptime history is disappointing. I know that they always explain what caused the downtime and I'm glad that they do. Explain what's wrong isn't enough, though. It seems like they have a poorly designed system that can be brought down by almost any piece failing. Their latest outage makes it sound like they don't know how to build a fault tolerant system at all. They have the parts there, like DRBD, but when things go wrong everything breaks anyway. It's like making backups but never testing them so that when you need them it turns out that they can't actually be used to restore anything.
It seems that they have the part that most companies miss, the communication, but that doesn't make up for the lack of reliability.
It seems that they have the part that most companies miss, the communication, but that doesn't make up for the lack of reliability.