Steel Industry ERP
Disaster Recovery for a 125 GB Odoo Production Host
Challenge
The problem
The client's largest Odoo production host carries the ERP for the whole business. It had daily backups, but no tested answer to the question of what happens if the machine itself is gone, and no second location.
Constraints
- This is the single most important machine the client runs. It cannot be experimented on
- Backups that live with the thing they are backing up are not backups. The copy had to leave the provider entirely
- A cold archive is not a recovery plan on its own. Restoring a database of this size from scratch under pressure takes longer than the business can wait
Approach
How we ran it
- Daily database and filestore backups pushed to object storage on separate infrastructure, so losing the provider does not lose the data
- A warm standby kept on a different provider in a different country, so the recovery path does not run through the same failure
- Monitoring on both, because a standby nobody checks is a standby that is quietly broken when it is finally needed
- Treated the production host as read-only for experiments. Anything that needed proving was proven elsewhere first
Ongoing managed service
By the numbers
- RAM production host under management
- 125 GBRAM production host under management
- off-site backup cadence
- Dailyoff-site backup cadence
- providers, in 2 countries
- 2providers, in 2 countries
Results
What we delivered
- Daily off-site backups to object storage, independent of the hosting provider
- Warm standby maintained on a different provider in a different country
- Recovery path that does not depend on the original machine or provider existing
Tech stack
Start a similar project →Odoo 13PostgreSQL 14Ubuntu 22.04Object StorageNginxUptime Kuma