Production infrastructure, CI/CD and observability
2 hardened servers, zero exposed app ports, tag-based rollback and pull backups.
Key figures
- production servers
- 2
- exposed application ports
- 0
- backup tiers
- 3
Context
The information system of a distribution and services SME relies on several critical applications: an API, an online store, an ERP and an identity provider. They must run continuously, deploy safely and stay recoverable after an incident.
My role
Design, setup, hardening and operation of the infrastructure, fully autonomously.
Solution
The infrastructure runs on 2 servers with separate roles: a cloud VPS for the API, the store and observability, and an on-premises server for the ERP and backups. Services sit in 5 isolated Docker networks and no application port is exposed: all traffic goes through Cloudflare Tunnel.
CI/CD builds images tagged by commit, so rolling back means redeploying a tag. Operations rely on Prometheus, Grafana, Loki and Alertmanager, with three-tier backups pulled in daily.
Key features
No exposed application ports
Applications are published through a tunnel, with no service listening directly on the internet.
Isolation and hardening
5 isolated Docker networks, hardened SSH, a firewall covering the DOCKER-USER chain, and fail2ban.
Tag-based rollback
Every image is tagged by commit, so going back to a stable version means redeploying its tag.
Full observability
Metrics, provisioned dashboards, centralised logs and alerting.
Engineering challenges
- 1
Pull-based backups
The backup server fetches the data itself, so a compromised server cannot erase the copy.
- 2
A silent failure uncovered
A backup had been failing for months: a pipe was masking its exit code.
- 3
Disk diagnosis with zero downtime
Disk usage brought down from 96% to 68% without service interruption, with an alert at 85%.
- 4
CI workflows that are tested too
actionlint and shellcheck check the workflows and shell scripts.
Tech stack
- Infrastructure
- Docker ComposeGitHub ActionsGHCRNginxLet's EncryptBash
- Security
- Cloudflare TunnelTailscalefail2ban
- Quality
- PrometheusGrafanaLokiPromtailAlertmanageractionlintshellcheck
- Data
- Borgrsync
A project of this scale?
Let's talk about your context: I'll tell you frankly what is feasible, and how long it takes.