Loading methodology...
Loading methodology...
Transparency
Every benchmark on Backup Arena is designed to be reproducible. This page documents the exact hardware, software, site profiles, timing rules, and verification checks so you can audit or replicate any result yourself.
Who runs this benchmark
Backup Arena is built and run by the team behind SafeGuard, a paid WordPress backup plugin that competes on this leaderboard. We are not a neutral third party and we don't claim to be.
SafeGuard gets no special treatment. It is entered on the same fixed snapshots, the same hardware, the same scope scoring, and the same three-iteration median as every other plugin. It does not automatically top every profile -- where a competitor is faster, that result stands and is published unchanged.
So don't take our word for it. Every snapshot is downloadable, every battle log is public, and the whole method below is reproducible. Rerun any result yourself and check our numbers.
Backup Arena benchmarks the premium/paid version of every plugin. This ensures each plugin is tested at its full capability — not limited by free-tier restrictions.
If a plugin has no paid version (the free version is the only version), we test that. The goal is to evaluate every plugin at its best.
| Plugin | Version Tested | Why |
|---|---|---|
| UpdraftPlus Premium | Paid | Premium features unlock full WP-CLI |
| Duplicator Pro | Paid | Faster engine, more options |
| Duplicator | Free | Free version is widely used |
| BackupBuddy | Paid | Paid-only plugin |
| WP Staging Pro | Paid | Pro backup features |
| SafeGuard | Paid | Full feature set |
| All-in-One WP Migration | Free | Uses custom adapter for CLI integration |
| BackWPup | Free | Free is the primary version |
| WPvivid | Free | Free core is the main product |
Every plugin runs the settings it ships with. The harness does not change plugin configuration, architecture, or compression, because a plugin we tuned would not be comparable to eight we did not. Compression in particular is left alone and measured instead: the method and the achieved ratio are recorded for every backup and shown on each result. That matters when reading the speed column, because a plugin that does not compress finishes sooner and writes a much larger archive, which is a real tradeoff rather than an advantage. Best Overall weighs time and size together for exactly that reason.
Before the three measured iterations begin, a single warm-up backup runs on each container. The warm-up result is discarded and never included in rankings.
The warm-up primes PHP opcache, MySQL's InnoDB buffer pool, and the Linux filesystem cache. Without it, the first iteration would be penalized by cold caches, skewing the median.
Every battle tests BOTH backup and restore. A plugin must successfully create a backup AND restore it for a hosting tier to count. Crashes are scoped to the tier they happen on: if a plugin runs out of memory, times out, or errors on one tier, it loses that tier, but the battle continues and the plugin is still measured on the tiers it survives. So a plugin that cannot back up a heavy site on 256 MB shared hosting can still post real times on a roomier VPS tier, and each result shows exactly where and why a crash happened.
Restore is tested by:
Not all backup plugins back up the same data. Some plugins (like SafeGuard, Duplicator, BackWPup) create full-site backups including WordPress core files and wp-config.php. Others (like All-in-One WP Migration) only back up wp-content and the database — by design, they're migration tools rather than disaster-recovery tools.
Each battle result displays backup scope badges showing exactly what each plugin included. This lets you make informed comparisons: a plugin that's faster because it backs up less isn't necessarily better.
Speed rankings use scope-adjusted times. A plugin that backs up less data receives a time penalty before the winner is decided, so backing up less is never free.
The penalty only covers bulk content: the database, uploads, plugins, themes and WordPress core. Skipping any of those saves real seconds and loses data that cannot be rebuilt, so seconds are a fair unit for the omission.
| Missing Category | Time Penalty |
|---|---|
| Database | +8s |
| Uploads | +6s |
| Plugins | +4s |
| Themes | +3s |
| WP Core | +3s |
| wp-config.php | not charged |
| .htaccess | not charged |
| Drop-ins (db, sunrise, object-cache, advanced-cache) | not charged |
For example, All-in-One WP Migration does not back up WordPress core files, so its raw time receives a +3s penalty before the speed winner is determined.
wp-config.php, .htaccess and the four drop-ins are detected and shown on every result, but they add no seconds to a plugin's time. Four of the nine plugins omit the two root files. Set against what those plugins actually cost, a charge of the size a completeness judgement would suggest looks like this:
| Plugin | Median total time | A 14s charge | As % of its runtime |
|---|---|---|---|
| All-in-One WP Migration | 4.2s | +14s | 331% |
| WP Staging Pro | 10.2s | +14s | 137% |
| Duplicator Pro | 23.4s | +14s | 60% |
| Duplicator | 25.9s | +14s | 54% |
Three separate facts rule that charge out:
Where a missing root file genuinely does break a restore, it is caught as the failure it is rather than priced as a time penalty: verification fails and that plugin loses the battle outright. Which components each archive contains is still detected and reported on every result page, as completeness rather than as seconds.
The drop-ins are a related but separate case. No site profile in the arena ships db.php, object-cache.php or sunrise.php, verified two ways (a scan of every snapshot, and the fact that no plugin has ever captured one in 180 battles). Charging for them would bill every plugin alike for files that were never there. They become chargeable only once a profile contains them.
Backup Arena scores plugins on whether they actually back up everything WordPress.org's official documentation says they should. We use a tiered penalty system based on what each missing component would cost a real user during recovery.
The table below is the severity ranking: what each missing component would cost a real user during recovery. It is what the completeness badges on every result page are based on. Only the bulk-content rows are also charged as time, for the reasons set out under root files and drop-ins.
| Component | Severity | Why it matters |
|---|---|---|
| Database | Critical - charged +8s | Without it, you have nothing — content, users, settings all gone |
| Uploads | Critical - charged +6s | Media library cannot be regenerated |
| Plugins | High - charged +4s | Plugins are user-installed and cannot all be re-downloaded |
| Themes | High - charged +3s | Themes can usually be re-downloaded but custom child themes cannot |
| WordPress core | Moderate - charged +3s | Officially recoverable from wordpress.org |
| wp-config.php | Reported only | Holds salts, custom defines and multisite topology. Plugins that omit it regenerate it on restore, and their restores verify — so it is shown as a completeness fact, not billed as time |
| .htaccess | Reported only | Custom security rules and redirects lost; default rules regenerate. Whether a plugin actually breaks permalinks is now checked directly after each restore |
| db.php drop-in | Reported only | Custom DB drivers (HyperDB, LudicrousDB) — site won't connect without it. No arena profile ships one |
| sunrise.php drop-in | Reported only | Multisite domain mapping — routing breaks without it. No arena profile ships one |
| object-cache.php drop-in | Reported only | Performance degradation only. No arena profile ships one |
| advanced-cache.php drop-in | Reported only | Performance degradation only. Present on one profile (News/Magazine) |
A crash is scoped to the hosting tier it happened on, not the whole battle. If a plugin runs out of memory on the shared tier, it loses that tier and the battle carries on, so it is still measured on the tiers it survives. That is a real result: a plugin that cannot finish inside a 256MB limit should lose to one that can.
If both plugins fail the same tier, nobody wins it and no rating changes.
A failure we cannot attribute to the plugin is not published as a loss. When a run shows a plugin failing the same phase on every single attempt, across every tier and every site, that is one repeated condition rather than many independent defeats, and the data cannot say whether the cause is the plugin or our own harness driving it. Those tiers are recorded as no-contest and flagged for review instead of counted. This is not hypothetical: a plugin once sat at 0-40 on this board because of a bug in our test harness, not in the plugin.
Infrastructure failures (Docker issues, network errors) that affect both plugins equally are excluded from rankings entirely.
Benchmarks run on dedicated-vCPU Hetzner Cloud servers that serve no other traffic, provisioned fresh for a run and destroyed afterwards, so no benchmark inherits another's state and no noisy neighbour shares the cores. Dedicated vCPU is the point: on shared vCPU the timings would depend on who else is on the host. A run uses several machines in parallel, each executing whole battles. Machine size follows the heaviest hosting tier being simulated — CCX23 (4 dedicated vCPU) for runs topping out at the 4-core tiers, CCX33 (8 dedicated vCPU) when the 8-core tiers are included.
| CPU | 8 vCPU (dedicated) |
| RAM | 32 GB ECC |
| Disk | 240 GB NVMe SSD |
| OS | Ubuntu 22.04 LTS |
| Network | Internal only (no public traffic) |
Each battle spins up four fresh containers: two WordPress + two MySQL instances, one pair per plugin. Containers run on an internal Docker network with no outbound internet access.
| WordPress image | wordpress:6.8-php8.2-apache |
| PHP version | 8.2 |
| MySQL image | mysql:8.0 |
| Docker network | --internal (no outbound internet) |
| WP container memory | 512 MB – 3 GB (by hosting profile) |
| PHP memory_limit | 256 MB – 1 GB (by hosting profile) |
| MySQL container memory | 1 – 2 GB (by hosting profile) |
| Container cleanup | Stopped + removed after every battle |
Each MySQL container uses a custom configuration file mounted at /etc/mysql/conf.d/bench.cnf:
[mysqld] innodb_buffer_pool_size = 1G max_allowed_packet = 256M innodb_flush_log_at_trx_commit = 0 tls_version=''
Disclosure: innodb_flush_log_at_trx_commit = 0
This is a deliberate deviation from MySQL defaults. It flushes the redo log once per second instead of on every transaction commit, improving write performance. This setting affects DB-heavy benchmark timing. It is applied identically to all plugin containers.
Every battle is tested across six hosting profiles that simulate real-world infrastructure constraints. Rather than benchmarking on a single generous server, this reveals how each plugin performs under the resource limits your host actually imposes.
| Profile | PHP Memory | Container RAM | CPUs |
|---|---|---|---|
| Shared | 256 MB | 512 MB | 1 |
| Managed WP | 512 MB | 1 GB | 2 |
| VPS Basic | 512 MB | 2 GB | 2 |
| VPS Standard | 512 MB | 1 GB | 4 |
| Dedicated | 1 GB | 2 GB | 4 |
| High Performance | 1 GB | 3 GB | 8 |
Each hosting profile runs 3 iterations per battle, producing 18 total iterations (6 profiles x 3 iterations). Resource limits are applied via Docker container constraints (CPU quotas, memory caps) so both plugins experience identical conditions.
Leaderboard rankings can be filtered by hosting profile, answering the question: "Which plugin wins on myhosting?" A plugin that dominates on a dedicated server may struggle on shared hosting with 256 MB of PHP memory.
Real websites do not sit idle while a backup runs. During benchmarks, simulated server traffic runs in the background to test how each plugin performs under realistic conditions.
Four types of background load are generated:
Load is applied at four intensity levels:
| Level | Request Rate | Simulates |
|---|---|---|
| None | 0 req/s | Idle server (baseline) |
| Light | 3 req/s | Low-traffic blog or portfolio |
| Medium | 10 req/s | Active business site |
| Heavy | 25 req/s | High-traffic store or news site |
The same load level is applied to both plugin containers in a battle, ensuring fairness. This answers the question: "How does this plugin perform when my site has active traffic?"
Eight pre-built WordPress sites cover the spectrum from a fresh install to enterprise scale. Each profile is a deterministic snapshot loaded identically for every battle. Sizes below are the measured contents of the snapshot, not estimates.
Five are currently in the published results (XS through M). The three largest are built and checksummed but have not yet been run across the full field, so no result on this site includes them. They are listed here so it is clear what the numbers do and do not cover.
Minimal WordPress blog — starter theme, 3 plugins (Akismet, Yoast SEO, Contact Form 7), ~50 posts with featured images.
Portfolio site — custom post types, code snippets plugin, image gallery, ~120 portfolio items with high-res images.
Small business site — page builder, forms, SEO, analytics, security, ~40 pages with media library of marketing assets.
News site — 2,000+ posts with categories/tags, multiple authors, ad integration, caching, image-heavy content. The only profile that ships an advanced-cache.php drop-in.
Mid-size WooCommerce store — 250 products, 570 orders, customer accounts, payment gateways, shipping integrations. Data generated using the WooCommerce PHP API (not bulk SQL): simple products (60%), variable products with Color/Size variations (30%), and grouped products (10%). Orders created via wc_create_order() with proper HPOS support. Includes coupons, product reviews, categories, and tags.
Large WooCommerce store — 2,000 products, 5,800 orders, subscriptions, product bundles, shipment tracking, multi-currency, advanced tax rules. Generated through the WooCommerce PHP API with HPOS support, same as the Standard profile.
Database-dominant store — a gigabyte of order and meta data against a comparatively small file tree, to separate database throughput from filesystem throughput.
Photography/video portfolio — 10,000+ high-res images, multiple image optimizers, gallery plugins, large media library. The inverse of the Enterprise profile: many files, modest database.
Each site profile is stored as a snapshot archive containing four components:
wp-content.tar.gz -- compressed uploads, themes, and plugin filesdatabase.sql.gz -- gzipped MySQL dumpmetadata.json -- profile name, file count, DB size, plugin listchecksums.json -- SHA-256 hash for every file in the snapshotBefore each battle, the Worker verifies the snapshot's SHA-256 checksum against the published value. The checksum for every profile is displayed in each battle's environment card and can be independently verified by downloading the snapshot.
Each battle runs three iterations. Before each iteration, both container pairs are reset to the clean snapshot state. Reset time (~10-20 seconds) is NOT included in timing -- the clock starts when the backup/restore command executes.
Each plugin's output is parsed into standard phases using a configurable phase parser:
| Phase | What it measures |
|---|---|
| SCAN | File and database discovery |
| DB_EXPORT | MySQL database dump |
| ARCHIVE | Compression of files into backup archive |
| RESTORE_FILES | Extraction and placement of backup files |
| RESTORE_DB | MySQL database import |
| VERIFY | Post-restore integrity checks |
Plugins whose output cannot be broken into phases receive a single BACKUP / RESTORE bar.
Different backup plugins use different archive formats, which affects both speed and output size:
| Format | Compression | Used by |
|---|---|---|
| ZIP | Yes (deflate) | UpdraftPlus, BackWPup, WPvivid, most plugins |
| TAR.GZ | Yes (gzip) | Duplicator Pro, some CLI-based plugins |
| WPRESS | No (proprietary) | All-in-One WP Migration |
| WPSTG | No (proprietary) | WP Staging |
Compression directly affects the fairness of speed comparisons. A plugin that skips compression will appear faster (less CPU work) but produces a larger archive. Conversely, a plugin that compresses well takes longer but saves storage and bandwidth.
To address this, Backup Arena shows both raw timing AND storage metrics for every battle: archive size with format badge, compression ratio (archive size vs. raw wp-content), and throughput in MB/s. The three-winner system (Speed, Storage, Best Overall) makes these trade-offs explicit -- a plugin that skips compression wins on speed but loses on storage.
Note on backup scope differences
All-in-One WP Migration backs up the database and the entire wp-content directory (plugins, themes, uploads, mu-plugins) but does not include WordPress core files (wp-admin/, wp-includes/) or wp-config.php — by design, it is a migration tool rather than a disaster-recovery tool. This smaller scope contributes to its faster backup times. Comparison pages show scope badges so you can judge speed differences in context.
Not all backup plugins back up the same data. Some include everything (database, plugins, themes, uploads, wp-config.php), while others skip certain categories. This makes speed comparisons unfair if one plugin simply backs up less data.
After each backup completes, Backup Arena inspects the backup output to detect which categories were included. It examines file listings, ZIP/TAR archive contents, and WPRESS file headers, looking for patterns like .sql files (database), plugins/, themes/, uploads/ directories, and wp-config.php.
Results display scope badges (DB, Plugins, Themes, Uploads, wp-config) for each plugin. When the two plugins have different scopes, an amber warning banner appears indicating that results may not be directly comparable. This gives readers the context they need to interpret timing differences fairly.
Each battle determines three winners:
Speed draws do not affect Elo ratings. Best Overall Elo only changes when one plugin Pareto-dominates the other; "Different Strengths" results count as draws with no Elo change.
The leaderboard number is an Elo rating -- the same system used for chess and competitive games. We use it because every battle is a head-to-head match and each plugin faces a different mix of opponents. A plain average time cannot account for that: beating a strong plugin should count for more than beating a weak one, and Elo is built to do exactly that.
Every plugin starts at 1000. After each battle the winner takes points from the loser, and how many points depends on who was expected to win. The expected score for plugin A against plugin B, and the resulting rating change, are:
E_A = 1 / (1 + 10^((R_B - R_A) / 400)) change = round(K * (actual - E_A)) K = 32, actual = 1 if win, 0 if loss
The K-factor of 32 caps how far a single battle can move a rating. Draws (speed margin under 0.5s) move nothing for either plugin.
Say an underdog rated 900 beats a favorite rated 1300. The favorite was heavily expected to win, so the upset is worth a lot:
E_underdog = 1 / (1 + 10^((1300 - 900) / 400)) = 1 / (1 + 10) = 0.091 Underdog wins: change = round(32 * (1 - 0.091)) = +29 -> 929 Favorite loses: change = round(32 * (0 - 0.909)) = -29 -> 1271
Flip it around: if that same 1300 favorite beats the 900 underdog, the win was expected, so it earns only round(32 * (1 - 0.909)) = +3. Upsets move the board; expected results barely nudge it. That is what keeps the ranking from simply rewarding whoever drew the weakest opponents.
A rating built on 5 battles is far less certain than one built on 200, so each rating carries a 95% confidence interval that shrinks as more battles are played:
CI = 1.96 * (400 / sqrt(n)) n = battles played
5 battles -> +/- 351 (provisional)
50 battles -> +/- 111
200 battles -> +/- 55Ratings with fewer than 5 battles are flagged Provisional. The interval is clamped to a minimum of +/-20 so the board never claims false precision.
Repeated failures from one cause count once. That formula assumes every battle is an independent observation, which stops being true when a single bug decides many of them. BackupBuddy failed to restore in 24 of its 40 managed-hosting battles, every one the same ActionScheduler class-redeclaration fault. That is one fact about the plugin observed 24 times, not 24 facts, and feeding it in as 24 makes the interval claim a precision the evidence does not support.
Failures are therefore grouped by harness version, hosting tier and phase, and each group counts as a single observation when sizing the interval. This is the cluster-sampling design effect at its most conservative. The rating itself is untouched— the losses are real and the plugin earned them — but the stated uncertainty widens to match. BackupBuddy's interval goes from +/-124 to +/-453, which is the honest reading of a record dominated by one bug.
The leaderboard ranks by speed. It does not tell you which plugin is best, and that is deliberate. Three things were measured rather than assumed.
1. Is a single rating even valid here? Elo assumes skill is transitive: if A beats B and B beats C, A should beat C. When that fails, ratings start depending on who happened to be scheduled against whom rather than on strength. Hamilton (2025) proves this and defines a statistic for it, the ratio of the cyclic to the transitive component of the advantage matrix under a Hodge decomposition, where below 1 means predominantly transitive. Computed across our pairwise tier medians:
speed ||transitive|| = 21.05 ||cyclic|| = 10.63 I(A) = 0.48 size ||transitive|| = 21.38 ||cyclic|| = 11.56 I(A) = 0.52
Both are comfortably below 1, so ranking each axis on its own is sound. The rating system is not the problem.
2. Do the axes agree with each other?No. Kendall's tau between the speed order and the archive-size order is -0.50. Not merely unrelated, actively opposed: across the whole field, the plugins that finish fastest write the biggest archives. That is the compression tradeoff showing up as a rule rather than as one plugin's quirk. Ranking by speed therefore ranks, in part, by who compresses least.
3. So who is actually best? Six of the nine plugins are Pareto-optimal, meaning no other plugin beats them on both speed and size at once. Only three are beaten on both. Any single “best” ranking would be picking a weighting between the two axes and applying it to you without saying so.
This is settled ground in benchmarking, not a novel position. SPEC published a global single-number indicator (SPECmark) and began strongly discouraging its use in 1991 after sustained criticism. Their Fair Use rules now encode the settlement: a composite is allowed as a derived valuebut may not be presented as the benchmark's own metric, the basis for comparison must be stated, and several benchmarks define required metrics that must be quoted alongside any headline figure. SPECpower does exactly that, pairing throughput with watts so neither travels alone. TechEmpower went the other way and introduced a composite score in Round 19; it remains the most contested part of their results.
So the arena reports both axes, marks which plugins are Pareto-optimal, and leaves the weighting to you. What we will not do is collapse a real tradeoff into one number and call it a verdict.
Restores land on a container that already has WordPress installed. A plugin that skips WP core therefore restores perfectly here, and would leave you with nothing on a bare server. All-in-One WP Migration is the clearest case: it is the fastest plugin in the arena and it is a migration tool, not a disaster-recovery tool. The +3s core penalty marks the difference; it does not capture the consequence.
Also unmeasured: cloud upload and download time (backups are written to local disk), incremental backup behaviour over time, multisite topologies, scheduled-job reliability, and support quality. Every number here is a single-run local backup and restore on a fixed snapshot.
Elo is never patched incrementally. After every new battle the entire history is replayed from the first result to the last, in order, so the board is always a pure function of the recorded battles -- no hidden state, no drift. The same replay produces separate ratings per site profile and per hosting profile, which is what the filters at the top of the leaderboard switch between.
Every restore is verified with four independent checks. If a plugin produces a corrupted restore, it is caught and shown publicly in the battle results.
File count match
Restored file count is compared against the original snapshot. Any discrepancy is flagged.
Sample checksum (SHA-256)
100 randomly-sampled files are hashed and compared against the original checksums.json. Catches silent corruption.
Database table check
SHOW TABLES count must match. Row counts for posts, postmeta, options, and users are compared against the original.
WordPress integrity
wp_options siteurl matches expected value. wp eval 'echo home_url();' returns the expected URL (WP-CLI check, no web server required).
Result: verified: true/false per plugin per iteration, displayed in every battle result.
In addition to file and database integrity checks, Backup Arena performs automated functional checks after each restore to verify the site actually works:
For WooCommerce profiles, additional checks are performed:
A failed verification is not treated as a crash, but it does decide the tier: the plugin that came back working takes it, however much slower it was. Both the per-battle result and the Elo replay apply the majority rule described under Winner Determination.
Backup Arena offers two execution modes. Users choose the mode when starting a battle, and the mode is displayed on every result page.
Both plugins run simultaneously on the same physical host. While each has dedicated MySQL and filesystem I/O via fully isolated container pairs (bench-wp-a + bench-db-a, bench-wp-b + bench-db-b), they share CPU cores and memory bandwidth.
A plugin's benchmark score may be marginally affected by the concurrent load of its opponent. This mirrors real-world hosting where WordPress shares resources with other processes.
To reduce positional bias, the container pair assignment is randomly swapped per iteration -- plugin A may run on container pair B and vice versa.
Each battle runs three iterations with a median calculation to reduce variance from transient resource contention.
Plugins run one at a time with a 3-second cache-settling pause between operations. The execution order is randomized per iteration (independent coin flip for backup and restore), eliminating cold/warm cache bias.
Sequential mode removes all resource contention between plugins, producing the fairest possible comparison at the cost of roughly doubling the total battle time.
Recommended when you need the highest-confidence results or when comparing plugins with very similar performance where shared-resource noise could affect the outcome.
Benchmark containers are sandboxed to prevent any plugin from affecting the host system or other battles.
--internal network with no outbound internet accessbackup_output_path is validated at three layers: frontend form, API endpoint, and Worker before every docker cp commandAll site profile snapshots are available for download. You can load them into the same Docker setup and run the benchmark scripts yourself to independently verify any result.
# Pull the benchmark containers docker compose -f docker-compose.bench.yml up bench-wp-a bench-db-a # Load a snapshot and run a backup docker exec bench-wp-a wp [plugin] backup ... # Compare your timings against published results
Questions about methodology? Contact us.