Configuration
Reference for backup-config.yaml — every key, default, and what changing it does.
The orchestrator reads two YAML files: backup-config.yaml (your preferences) and the cluster's prod-config.yaml (referenced by the prod_config: key, gives SSH details for the 3 production nodes). The example file at automation/backups/backup-config.example.yaml is the source of truth for the schema.
Top-level keys
prod_config
Path to the production prod-config.yaml. Relative paths resolve against the automation/backups/ directory. Default: ../production/prod-config.yaml.
backup_* (host details)
backup_private_ip, backup_ssh_host, backup_ssh_user, backup_ssh_key. If backup_ssh_host is blank, falls back to backup_private_ip. When aws-provision runs with backup_node.enabled: true, these are written into provision-output.yaml automatically.
backup_repo_root
Where every repo lives on the backup host. Default /var/lib/openg2p-backup — matches the cloud-init mount point of the dedicated EBS data volume. Change this only if you want repos somewhere else (e.g. a custom block storage path on a non-AWS install).
Encryption — passphrase files
restic_passphrase_file: "~/.openg2p/keystore/restic.pass"
pgbackrest_passphrase_file: "~/.openg2p/keystore/pgbackrest.pass"
etcd_at_rest_key_file: "~/.openg2p/keystore/etcd-aescbc.key"~ is expanded to the operator's home. If a file is empty/missing at install time, the orchestrator generates a random passphrase and writes it back. You must then move that file into your p12 keystore — the automation prints a reminder.
Group toggles
Each group is independently switchable. Disabling a group:
Skips its install steps
Comments out its cron entries on the backup host
Causes
run/verify/drillto skip itReports it as
disabledinstatusoutput
objectstore defaults to false when the key is missing (unlike the other groups, which default to on). Enable it only after rclone credentials and a restic password file are in place — see Object store below.
You can disable a group later by editing backup-config.yaml and re-running install — the orchestrator regenerates the cron file.
Retention
Defaults give roughly 6 months of granular history. Government data-retention policies may require more — increase keep_monthly (and ensure the data volume can hold it). The drill (drill subcommand) does not touch retention.
Schedules
Standard cron syntax. The orchestrator renders these into /etc/cron.d/openg2p-backup on the backup host at install time. Edit the cron file directly to test changes; commit them back to backup-config.yaml so the next install re-applies them.
The defaults stagger: PG at 02:00, etcd every 6h offset by 15 minutes, rancher and NFS at 03:00–03:30 (rancher writes a tarball that NFS then captures, so order matters), objectstore at 04:00 when enabled, drill on Sunday at 05:00 after Saturday's runs, daily email at 07:00, WAL probe every 5 minutes.
About schedules.rancher: nightly rancher backups are driven by the in-cluster Schedule CR (manifests/rancher-backup-schedule.yaml), not by a cron entry on the backup host. The schedules.rancher value is informational only today — it documents the intended cadence. Ad-hoc rancher backups can still be triggered from the laptop with ./openg2p-backup.sh run --component rancher.
PostgreSQL
canary_table: optional. If set, the weekly drill runs SELECT count(*) FROM <canary_table> against the temporarily-restored Postgres to confirm the backup is application-readable, not just byte-readable. Pick a table that is always present and has predictable content (e.g. a users table with a known minimum row count).
archive_timeout_seconds: how long Postgres waits before forcing a WAL switch. Lower = better RPO but more WAL files. 60s is a good default; under 30s starts wasting space.
NFS data
Use paths as an explicit allowlist when you don't want the entire export. The orchestrator's bash YAML parser doesn't natively handle nested arrays — if you need multiple paths, use nfs.path1, nfs.path2 flat keys (the parser handles those):
System logs are excluded by default because OpenSearch already retains them.
ResourceSet
The live ResourceSet CR applied at install is automation/backups/manifests/rancher-backup-resourceset.yaml in the deployment repo. Edit that manifest to change what is captured. At install time, the orchestrator validates each apiVersion entry against the live cluster (kubectl api-resources) and warns about any unknown API group — it does not fail install if optional operators (cert-manager, Istio, Keycloak, etc.) are not yet deployed.
run --component rancher also re-applies this ResourceSet before creating an ad-hoc Backup CR. That matters after Rancher / rancher-backup-crd upgrades, which commonly drop custom ResourceSets and leave backups failing with resourcesets.resources.cattle.io "openg2p-resource-set" not found.
The operator CRD uses strict decoding:
There is no top-level
namespaceRegexpor booleancontrollerReferencesfield.Namespace scoping is per
resourceSelectorentry (namespaces/namespaceRegexp, Go RE2 only — no negative lookahead).The shipped manifest captures all namespaces for DR completeness; system-namespace objects are small and mostly helmfile-recreatable.
The resource_set: block in backup-config.example.yaml documents the intended policy for operators customizing the manifest — it is not rendered into the CR at install time.
Example selectors (see the manifest for the full list):
Add custom CRD groups in the manifest if your environment installs additional operators (e.g. eventing.knative.dev).
rancher-backup operator storage
The backup-restore-operator writes encrypted tarballs to a PVC mounted at /var/lib/backups inside the operator pod. Do not set storageLocation on Backup / Restore CRs — the operator only supports explicit storageLocation for S3; PVC storage is configured at the Helm chart level.
Static NFS PV (stable location). The chart's default dynamic persistence names the PVC <release>-<helm-revision>, so every helm upgrade provisions a fresh empty volume and strands prior backups. Install therefore:
Reads the NFS
server+sharefrom the cluster'spvc_storage_classStorageClass.Creates
/srv/nfs/<cluster>/rancher-backupon the storage node (mode0777so the operator pod can write).Applies a static NFS
PersistentVolumenamedopeng2p-rancher-backup-store.Installs the chart with
persistence.enabled=true,persistence.storageClass=-(no dynamic provisioning), andpersistence.volumeName=openg2p-rancher-backup-store.
Tarballs land at /srv/nfs/<cluster>/rancher-backup/ on the NFS export (captured downstream by the nfs restic job). With encryptionConfigSecretName set, files are named *.tar.gz.enc.
To re-provision rancher storage after a chart/config change without re-running the full install (which restarts RKE2 for etcd):
Etcd encryption-at-rest
Default: disabled. To turn it on, run ./openg2p-backup.sh install --enable-secret-encryption during a maintenance window. This is a separate, deliberate action because it restarts the apiserver. The flag overrides this config key only for that run; flipping the key to true does not automatically enable encryption — you must pass the CLI flag.
Tool versions
Pinned to known-good versions. Bump after testing in a non-production environment. Choose a rancher_backup_chart whose rancher-version / kube-version annotations match your cluster (OpenG2P production default: Rancher 2.15.1, RKE2 v1.35.8+rke2r1):
After a Rancher or BRO chart upgrade, re-run:
That refreshes CRDs + operator, re-applies the ResourceSet / Schedule, and tolerates a failed rancher-backup-patch-sa hook (manual SA patch + --no-hooks retry).
Monitoring
When enabled: true, each successful/failed run writes Prometheus textfile metrics under $backup_repo_root/metrics/. install applies manifests/prometheusrule-backup.yaml.template into monitoring.namespace (default cattle-monitoring-system). If the CRD or namespace is missing, install warns and continues — re-run after Rancher Monitoring is up.
Prometheus only fires those rules if it scrapes the backup host metrics (node_exporter textfile collector or Pushgateway). Details and how to test: Alerting.
Alerting (email)
Copy automation/backups/roles/backup-host/smtp.env.example to the path in smtp_env_file, fill SMTP settings, set email_enabled: true, and re-run install. That copies the file to /etc/openg2p-backup/smtp.env (mode 0600) on the backup host.
daily-report and failure notifications use Python smtplib on the backup host. For the in-cluster OpenG2P mail chart (openg2p/smtp), the Service is ClusterIP:25 and typically only relays from the pod CIDR — the backup host on the VPC usually needs a NodePort (or equivalent) plus RELAY_NETWORKS including 172.29.0.0/16. See Alerting.
Object store (MinIO/S3)
Opt-in. Installs rclone + restic on the backup host, mounts remote:bucket read-only, and takes a restic snapshot into $backup_repo_root/objectstore-restic. Repositories and credentials:
/etc/openg2p-backup/rclone.conf/etc/openg2p-backup/restic-objectstore.env(RESTIC_PASSWORD=…)
Last updated
Was this helpful?