← Articles
Database Reliability

Database Backups: What Companies Get Wrong

Backup success messages can conceal serious recovery gaps. Companies need recovery objectives, complete backup chains, protected copies, monitored jobs, and regular restores that prove both the data and the runbook.

Most companies with SQL Server have backup jobs. Far fewer have evidence that they can recover the business service within an acceptable time and to an acceptable point. The gap appears during ransomware, corruption, storage failure, or an accidental data change, when a green job history is mistaken for a tested recovery capability.

Starting with a schedule instead of a recovery objective

A nightly full backup is a schedule, not a requirement. The business must first define the maximum tolerable data loss and the time available to restore service. Those are commonly described as the recovery point objective and recovery time objective. They should be stated for each important system because payroll, an internal archive, and live order processing rarely have identical needs.

If the business can lose no more than fifteen minutes of committed orders, a nightly full backup cannot meet the requirement. A transaction log backup strategy may be necessary. If service must return quickly, restore duration, data transfer, validation, and application startup all belong in the calculation.

Trusting job success without checking the artifact

A job can report success while writing to a location with inadequate retention or while omitting a newly created database. Files can later be deleted, truncated, or made inaccessible by a credential change. Monitor both execution and outcomes: expected databases, backup type, finish time, age, size anomalies, checksum settings, destination capacity, and copy or replication status.

Use backup verification as one layer, but understand its limits. Reading backup metadata and checksums can detect some damage; it does not run database recovery or prove the restored application works. Only a restore exercises the relevant sequence.

Breaking the log chain unintentionally

Point-in-time recovery under the full recovery model depends on an unbroken transaction log backup chain after an appropriate full backup. Switching recovery models, failing log backups, replacing a database, or misunderstanding copy-only operations can change what recovery is possible. Log files that grow without limit often indicate that log backups are absent or failing, but truncation state must be diagnosed rather than guessed.

Document which full, differential, and log backups form the recovery sequence. During an incident, restore operators should not have to infer file order from names alone. Keep authoritative metadata and a tested script that can select the intended recovery point.

Keeping every copy in one failure domain

A backup stored on the same server protects against some logical failures but not loss of that server. A copy on a share joined to the same identity boundary may still be vulnerable to the same destructive account or ransomware event. Recovery design should include separate failure domains and, where appropriate, an immutable or offline copy.

  • Encrypt backups when required and protect the keys or certificates separately.
  • Restrict write and delete rights more tightly than read access needed for restore.
  • Transfer copies over protected channels and monitor arrival at the secondary location.
  • Set retention from business, regulatory, and corruption-discovery needs.
  • Test restoration when the primary identity or network service is unavailable.

Encryption without key recovery is permanent data loss. Record key custody and test it with the same discipline as the backup file.

Never timing a restore

Restore tests often stop when the database reaches an online state. A complete exercise measures file retrieval, restore, recovery, integrity validation, login and permission repair, job enablement, connection changes, and application verification. Large databases should be tested on infrastructure representative of the recovery target because storage throughput can dominate the timeline.

Evidence from a useful restore test
AreaRecorded result
Recovery pointExact UTC time reached and expected data-loss window
Recovery timeDuration by retrieval, restore, validation, and service startup
IntegrityDatabase checks and business reconciliation completed
AccessRequired logins, jobs, keys, and application connections verified

Ignoring human access and ownership

A technically valid backup is useless if no available person can reach it during an outage. Avoid dependence on one employee's account, memory, or laptop. Name primary and backup owners, protect credentials in an approved secrets system, and make the recovery runbook available during a network or identity failure.

The runbook should use exact server roles and approved locations without embedding passwords. It should explain who can authorize a restore, how to prevent users from connecting too early, how to choose a recovery point, and how the business validates results.

Treating high availability as a backup

Availability replicas and failover services reduce downtime for certain failures. They can also reproduce accidental deletes, damaging updates, or corruption. They usually do not provide long historical retention. High availability and backups solve related but different problems, and a mature design uses both according to defined scenarios.

A backup strategy is credible only when a dated restore record proves that people, files, credentials, infrastructure, and instructions worked together.

Review recovery objectives at least when applications, volumes, infrastructure, or business obligations change. Automate routine validation where practical, then conduct periodic full exercises with the actual owners. The work is complete when the organization can state what it can recover, to when, in how long, and show the evidence.

Working through a similar problem?

Discuss your database question.

Start a conversation