# Benchmarks and pricing

The [benchmark page](https://imgt.fashn.ai/benchmarks) provides live latency
charts, smaller local allocations, and an interactive monthly cost model.
Measurements and rates were checked on October 5, 2026.

## Live comparison

Same 50 PNG sources, width 640, quality 98 and WebP effort 4. One client entered
through São Paulo edge locations; Cloudflare containers were in North America.
These were sequential single runs with warm containers. Complete-response latency
ends after the response body, before output validation.

| Workload, p95 | Four 2-CPU containers | Eight 1-CPU containers | Vercel |
| --- | --- | --- | --- |
| 20 fresh images/s, 1,200 requests | 1.148 s | 1.462 s | 1.341 s |
| 50 fresh images together | 1.755 s | 2.637 s | 1.377 s |
| 200 fresh images together | 6.525 s | 5.636 s | 1.508 s |
| 200 cached images together | 230 ms | 222 ms | 193 ms |

All requests succeeded. All 1,400 comparable fresh outputs matched Vercel byte
for byte. A separate three-minute run completed 3,600 fresh transforms at 20/s,
without errors, at a 1.110 s p95. Four-shard fresh-container warmups took up to
7.376 s and eight-shard warmups up to 9.335 s. Cold provisioning can exceed the
20-second deadline. These results do not establish a global latency SLO.

## Memory and allocation

The measured renderer does not require the full provisioned allocation. A local
100-request large-input test, including a 50-megapixel RGBA image and inputs padded
to 25 MiB, peaked at 628 MiB per container. Cloudflare custom allocations require
at least 3 GiB per CPU, so a 2-CPU container reserves at least 6 GiB. See
[resource limits](https://developers.cloudflare.com/containers/platform/limits/).

Reducing shard count or CPU capacity reduces provisioned memory charges. It also
reduces uncached burst capacity. Four 2-CPU containers completed the 20/s comparison. The current default uses two
2-CPU containers, with two encoders per container. Smaller configurations are
experiments, not general production recommendations. A live two-1-CPU-container repeat
completed 300/300 fresh transforms at 5/s with a 1.222 s p95. Its 50-image burst
was 4.201 s and its 200-image burst 15.254 s. At 10/s it completed 598/600, with
two client timeouts. The initial 1/s attempt also had an upstream error and two
client timeouts; the repeated 1/s run completed 30/30 with a 5.380 s p95. Timeout
causes remain unresolved. All 1,178 successful fresh repeat outputs matched
Vercel, and cache and conditional 304 checks passed. Detailed smaller-allocation
results are recorded in [performance notes](performance.md).

### Two 2-CPU containers

A later live trial allocated 12 GiB total RAM and retained two encoders per
container. Warm 5/s completed 300/300 at 1.047 s p95; warm 10/s completed
600/600 at 1.084 s. Fresh 50- and 200-image bursts completed every request at
2.996 s and 9.812 s respectively. Two idle-stop/restart cycles completed 50/50
fresh images each at 3.515 s and 3.769 s, with new boots for both containers.
Cached repeats completed 200/200 at 280 ms and conditional validation passed.

The first provisioning burst immediately after deployment returned 50 HTTP 500
responses. Its cause remains unresolved. The five-minute 10/s confirmation then
completed 2,998/3,000, with two 30-second client timeouts. Successful responses
had a 1.094 s p95 and 28.885 s maximum. Admission p95 was zero and encoder
queue p95 was 45.9 ms; the timeouts do not establish a CPU capacity limit.
All 4,282 successful fresh outputs across the warm trial, restart cycles and
longer confirmation matched Vercel. That trial restored the four-container allocation
because the smaller configuration failed its reliability check. See the complete
[performance notes](performance.md) for failures and response maxima.

### Reliability repeat

A repeat used a 60-second client limit with the unchanged 20-second transform
deadline. First provisioning returned 50/50 fresh images at 8.862 s p95.
A five-minute 3/s phase completed 900/900 at 1.105 s; the five-minute 10/s
phase completed 2,999/3,000 at 1.099 s. There were no timeouts. The one failure
was a source HTTP 500, surfaced as HTTP 502 in 440 ms; repeating it later
succeeded in 799 ms. Fresh 50- and 200-image bursts returned every image at
2.988 s and 9.208 s p95, cached repeats at 223 ms. An idle restart returned
50/50 at 4.131 s. All 4,252 successful fresh outputs matched Vercel.

Every successful response arrived within the earlier 30-second limit; the
longer client timeout did not recover a slower response in this repeat.
Earlier failures remain unexplained. A subsequent source-retry rollout caught
explicit Durable Object reset errors: the 10/s test stopped at 414 requests,
with 408 successes and six HTTP 500 responses. A later 200-image burst
returned 199 errors; repeating it returned all 200, including 193 retained in
R2. Those failures were rollout interruptions, rather than encoder deadline
failures. imgt now retries the source GET once on temporary failures and
reacquires a Durable Object stub up to twice on exceptions marked retryable, excluding
overloaded objects. Container uploads and HTTP errors are not replayed.
See [request recovery behavior](architecture.md#slow-requests-and-errors).

### Rollout and health checks

Later startup checks retained two 2-CPU containers but increased each transform
attempt to 60 seconds and the client limit to 90 seconds. They did not pass:

| Check | Successful images | Other results |
| --- | --- | --- |
| One reset retry, earlier 20 s deadline | 38/50 | Eleven HTTP 503, one HTTP 500 |
| 60 s deadline, settings before health | 0/50 accepted | 39 HTTP 504, four HTTP 503, seven version-check failures |
| Settings moved to background restoration | 0/50 | 49 HTTP 504, one HTTP 503 |
| Two-second health probes | 33/50 | Sixteen HTTP 504, one version-check failure |

The last check's successful p95 was 44.515 s and maximum 44.908 s. One shard
became healthy in 40.204 s; the other exhausted its 60-second readiness limit.
A subsequent two-shard warm-up completed 1/2 and stopped before sustained load.
These results do not establish two-container reliability or a CPU capacity limit.
The clean confirmation below used fresh runtime-v11 shard IDs with the same allocation and
renderer; stored derivatives remain compatible.

Cloudflare reported a [regional Durable Objects and Containers incident](https://www.cloudflarestatus.com/incidents/z28czd65h29l)
from October 5, 17:14 to 22:35 UTC in Eastern North America, matching the container
region. Earlier trials overlapped it; the latest checks followed its reported
resolution. This is relevant context, not an established cause of every failure.
The new pending-I/O compatibility flag was already enabled in the live Worker.

### Clean confirmation with fresh shards

A fresh runtime-v11 pair completed the immediate 50-image provisioning burst
and all 3,000 arrivals at 10/s over five minutes. It retained two encoders per
container, 12 GiB total RAM, a 60-second job deadline and 90-second client limit.

| Workload | Success | Complete-response p95 | Maximum |
| --- | --- | --- | --- |
| First provisioning | 50/50 | 12.394 s | 21.484 s |
| 10/s, five-minute arrival window | 3,000/3,000 | 1.571 s | 67.690 s |
| 50 fresh images together | 50/50 | 2.733 s | 2.913 s |
| 200 fresh images together | 200/200 | 10.330 s | 26.125 s |
| 200 cached images together | 200/200 | 296 ms | 533 ms |
| 50 fresh images after idle stop | 50/50 | 4.237 s | 12.668 s |

Both containers stopped after idle time and restarted with new boot IDs.
All 3,352 fresh outputs, including the restart gallery, matched Vercel.
All 20 responses slower than 20 seconds were rechecked from edge cache in at most 70 ms. The sustained phase took
321.558 seconds including delivery after the arrival window. Queueing remained
short, but delivery outliers remain unresolved: the 67.690-second response had
headers after 740 ms and a 190 ms transform job. A job deadline does not bound
full-body delivery. The two-container allocation remains the smaller starting
point for this tested workload; its 20/s capacity remains unverified. These finite
results do not establish Vercel-equivalent tail latency or long-term reliability.
The earlier failures remain visible above.

## Cost model

The public calculator uses a synthetic scenario: two million fresh derivatives,
three million total requests, 120 KiB per output, eight million billed Vercel read
units, 0.14 consumed CPU seconds per fresh transform, and every container awake
for a 30-day month. It assumes Cloudflare account allowances are available.

| Allocation | Total reserved RAM | Modeled imgt monthly subtotal |
| --- | --- | --- |
| One 1-CPU container | 3 GiB | $38.02 |
| Two 1-CPU containers | 6 GiB | $70.32 |
| Two 2-CPU containers | 12 GiB | $109.20 |
| Four 2-CPU containers | 24 GiB | $187.69 |
| Eight 1-CPU containers | 24 GiB | $214.14 |

The equivalent modeled Vercel image meters are $223.20 at iad1 on-demand rates,
or $359.52 at gru1 rates, before plan credits and delivery charges. These are
estimates, not bills or savings guarantees. The smaller profiles provide less
uncached throughput and have not established live parity with the larger profile.

imgt subtotals include the $5 Workers plan, provisioned container memory and disk,
consumed CPU, modeled Durable Object duration/requests, R2 Standard storage and
operations, Worker requests/CPU, and North American container output transfer.
The model conservatively sends every request through the Durable Object and R2,
adds an extra R2 read per transform, and retains 30 days of unique derivatives.
Workers bill both the gateway and cached loopback invocation, with an assumed
combined 7 ms CPU per external request; Durable Objects stay active for the
container's entire awake period at 128 MB each. R2 and Durable Object excess
billing units are rounded up. These assumptions can differ from actual usage.

CPU measurements were approximately 0.12 to 0.15 seconds per transform locally.
The default rounds to 0.14. Provisioned CPUs are not billed as active continuously.
Awake time cannot be below the modeled consumed CPU requirement; allocations with
insufficient monthly CPU capacity are flagged. Uniform traffic and the two-minute
idle timeout can keep all containers awake even at modest average throughput.

Vercel write units are fresh transforms multiplied by output size rounded up to
8 KiB. Read units must come from billed usage, not image request counts; recently
cached regional reads may avoid the read meter. Two million transformations over
30 days average 0.77/s, but monthly counts cannot establish peak traffic.

Excluded: Vercel subscription, credits, CDN requests and data transfer;
source-host transfer on either provider; excess logs/traces, Durable Object storage,
and container idle/startup CPU overhead. Taxes and support time are excluded.
Shared allowances, caching, retained unique objects and geography affect costs.
Use a private `imgt` bucket with an enabled `derivatives/` lifecycle rule deleting
after 30 days, as described in [deployment](deployment.md#storage-cleanup).

Price sources:

- [Cloudflare Containers](https://developers.cloudflare.com/containers/platform/pricing/)
- [Cloudflare Durable Objects](https://developers.cloudflare.com/durable-objects/platform/pricing/)
- [Cloudflare Workers](https://developers.cloudflare.com/workers/platform/pricing/)
- [Workers Cache invocation billing](https://developers.cloudflare.com/workers/cache/#pricing)
- [Cloudflare R2](https://developers.cloudflare.com/r2/pricing/)
- [Vercel image billing units](https://vercel.com/docs/image-optimization/limits-and-pricing)
- [Vercel iad1 rates](https://vercel.com/docs/pricing/regional-pricing/iad1)
- [Vercel gru1 rates](https://vercel.com/docs/pricing/regional-pricing/gru1)

Raw inputs, credentials, per-request logs and research scripts are excluded from
the public site and normal repository checkout.
