All projects
DevOps

Production infrastructure, CI/CD and observability

2 hardened servers, zero exposed app ports, tag-based rollback and pull backups.

Key figures

production servers
2
exposed application ports
0
backup tiers
3

Context

The information system of a distribution and services SME relies on several critical applications: an API, an online store, an ERP and an identity provider. They must run continuously, deploy safely and stay recoverable after an incident.

My role

Design, setup, hardening and operation of the infrastructure, fully autonomously.

Solution

The infrastructure runs on 2 servers with separate roles: a cloud VPS for the API, the store and observability, and an on-premises server for the ERP and backups. Services sit in 5 isolated Docker networks and no application port is exposed: all traffic goes through Cloudflare Tunnel.

CI/CD builds images tagged by commit, so rolling back means redeploying a tag. Operations rely on Prometheus, Grafana, Loki and Alertmanager, with three-tier backups pulled in daily.

Key features

  • No exposed application ports

    Applications are published through a tunnel, with no service listening directly on the internet.

  • Isolation and hardening

    5 isolated Docker networks, hardened SSH, a firewall covering the DOCKER-USER chain, and fail2ban.

  • Tag-based rollback

    Every image is tagged by commit, so going back to a stable version means redeploying its tag.

  • Full observability

    Metrics, provisioned dashboards, centralised logs and alerting.

Engineering challenges

  1. 1

    Pull-based backups

    The backup server fetches the data itself, so a compromised server cannot erase the copy.

  2. 2

    A silent failure uncovered

    A backup had been failing for months: a pipe was masking its exit code.

  3. 3

    Disk diagnosis with zero downtime

    Disk usage brought down from 96% to 68% without service interruption, with an alert at 85%.

  4. 4

    CI workflows that are tested too

    actionlint and shellcheck check the workflows and shell scripts.

Tech stack

Infrastructure
Docker ComposeGitHub ActionsGHCRNginxLet's EncryptBash
Security
Cloudflare TunnelTailscalefail2ban
Quality
PrometheusGrafanaLokiPromtailAlertmanageractionlintshellcheck
Data
Borgrsync

A project of this scale?

Let's talk about your context: I'll tell you frankly what is feasible, and how long it takes.