MEASURED OCTOBER 5, 2026

Performance has
a price.

Measured latency. Smaller allocations. Monthly cost estimates.
Choose the capacity your images actually need.

1.15 sLive p95 at 20 fresh images/s
0 / 3,600Errors in the three-minute confirmation
628 MiBPeak per container in a local large-input test

01 / LIVE LATENCY

Compare complete response time.

Same sources. 640 px. Quality 98.
WebP effort 4. Warm containers.

p95 · lower is faster
imgt · 4 × 2 CPUs1.755 s
imgt · 8 × 1 CPU2.637 s
imgt · 2 × 1 CPU4.201 s
imgt · 2 × 2 CPUs2.996 s
Vercel1.377 s

Live requests from one client through São Paulo edge locations. Cloudflare containers ran in North America. Each comparison ran once. The two-container trials ran later; see their failures and variable delivery latency below. These results are not a global latency guarantee.

imgt was close to Vercel for sustained requests and cached delivery. Vercel was faster for large bursts of uncached images. All fresh outputs in the matched 1,400-request comparison were byte-identical.

All live results and sample counts
Complete-response p95; every request succeeded
Workload 4 × 2 CPUs 8 × 1 CPU Vercel
20/s, 1,200 fresh requests 1.148 s 1.462 s 1.341 s
50 fresh requests together 1.755 s 2.637 s 1.377 s
200 fresh requests together 6.525 s 5.636 s 1.508 s
200 cached requests together 230 ms 222 ms 193 ms

Four 2-CPU containers also completed 3,600 fresh requests at 20/s with zero errors and a 1.110 s p95. Fresh-container warmups took up to 7.376 s for four shards and 9.335 s for eight. Cold starts are excluded from the charts and can exceed the 20 s request deadline.

02 / SMALLER ALLOCATIONS

How small can you go?

Memory follows the CPU allocation.
Your traffic determines the CPU need.

The renderer does not need 6 GiB to process the tested images. Cloudflare custom containers require at least 3 GiB per CPU, so each 2-CPU container reserves 6 GiB. Reducing the number of containers or CPUs reduces that memory bill.

Two 2-CPU containers, live

12 GiB total reserved RAM
Warm delivery and idle restarts; complete-response p95
Workload Success p95
5/s, 300 fresh requests 300/300 1.047 s
10/s, 600 fresh requests 600/600 1.084 s
10/s, five minutes, 3,000 fresh requests 2,998/3,000 1.094 s
50 fresh images together 50/50 2.996 s
200 fresh images together 200/200 9.812 s
200 cached images together 200/200 280 ms
50 images after idle stop, first cycle 50/50 3.515 s
50 images after idle stop, second cycle 50/50 3.769 s

The first 50-image provisioning attempt immediately after deployment returned 50 HTTP 500 responses. The later warm run and both idle restarts succeeded. The first failure remains unresolved and is retained in the results. The five-minute run had two 30 s client timeouts; their causes remain unresolved. Successful warm fresh outputs and the 100 restart outputs matched Vercel; cache and 304 checks passed. Warm delivery maxima were 17.486 s at 10/s and 28.644 s in the 200-image burst. The second restart's maximum was 19.157 s. The five-minute run's maximum successful response was 28.885 s. P95 does not describe these slower responses. The deployment retains four 2-CPU containers because the smaller trial failed its reliability check. The measured encoder queue was short, so these timeouts do not establish a CPU capacity limit.

Two 1-CPU containers, live

6 GiB total reserved RAM
Warm repeat trial; p95 of successful responses
Workload Success p95
1/s, 30 fresh requests 30/30 5.380 s
5/s, 300 fresh requests 300/300 1.222 s
10/s, 600 fresh requests 598/600 1.510 s
50 fresh images together 50/50 4.201 s
200 fresh images together 200/200 15.254 s
200 cached images together 200/200 467 ms

The first 1/s trial returned 27/30: one upstream HTTP 500 and two 30 s client timeouts. The repeat had two client timeouts at 10/s. Long low-load client waits were observed outside the instrumented transform job: one 18.4 s delivery reported a 385 ms job and no admission/encoder queue. The timeout causes remain unresolved. These observations do not establish reliable 10/s capacity. All 1,178 successful fresh repeat outputs matched Vercel; cached responses and conditional 304 validation passed.

Less reserved memory buys a smaller encoding capacity. The two-container live trial was slower for fresh bursts, while its 5/s run completed without errors. Keep reliability and your uncached latency target alongside the cost estimate.

Warm local replay

640 px · quality 98 · effort 4
Successful requests / attempted; p95 of successful requests
Allocation / total RAM 1/s · 30 requests 5/s · 150 10/s · 300 20/s · 600 50-image burst
1 × 1 CPU / 3 GiB 30/30; 0.243 s 150/150; 0.212 s 300/300; 7.712 s 356/600; 16.012 s 50/50; 6.222 s
2 × 1 CPU / 6 GiB 30/30; 0.246 s 150/150; 0.216 s 300/300; 0.385 s 600/600; 9.612 s 50/50; 3.443 s
4 × ¼ CPU / 4 GiB 30/30; 1.075 s 150/150; 2.179 s 298/300; 15.191 s 309/600; 19.545 s 50/50; 7.500 s

Two CPUs completed the 20/s arrival window, but their queue grew and needed another 10 seconds to drain. That is overload, not sustained 20/s capacity. One CPU rejected 244 of 600 requests at 20/s. The fractional tier had deadline and admission failures at 10/s and 20/s. Cloudflare lists basic in its limits/pricing tables, but the native start API documentation and current types omit it; native deployment support is unverified.

The table above measures local capacity, rather than deployed delivery latency. Local sources and simulated R2 writes remove network and provisioning costs. More CPU improves uncached burst latency; cache hits skip encoding entirely.

03 / COST MODEL

See what you would pay.

USD · 30-day month.
Estimates, not measured bills.

imgt can cost more at low traffic if containers stay awake. At higher volume, smaller allocations can be substantially cheaper. Memory and disk are billed while allocated; CPU is billed for active usage. Change the assumptions below to compare.

Modeled monthly cost

lower is cheaper
imgt · 1 × 1 CPU$38.02
imgt · 2 × 1 CPU$70.32
imgt · 2 × 2 CPUs$109.20
imgt · 4 × 2 CPUs$187.69
imgt · 8 × 1 CPU$214.14
Vercel$223.20

imgt bars include the $5 Workers plan, container memory, active CPU, disk, Durable Objects, R2, Worker usage and modeled container output transfer. Vercel bars cover image transformations and cache meters only.

Cost breakdown by allocation
Modeled imgt monthly charges in USD
Component 1 × 1 CPU 2 × 1 CPU 2 × 2 CPUs 4 × 2 CPUs 8 × 1 CPU
Workers plan$5.00$5.00$5.00$5.00$5.00
Container memory$19.21$38.66$77.54$155.30$155.30
Active container CPU$5.15$5.15$5.15$5.15$5.15
Container disk$0.31$0.68$0.68$1.40$2.85
Durable Objects$0.30$12.80$12.80$12.80$37.80
R2 storage and operations$8.04$8.04$8.04$8.04$8.04
Workers usage$0.00$0.00$0.00$0.00$0.00
Container output transfer$0.00$0.00$0.00$0.00$0.00
Modeled subtotal$38.02$70.32$109.20$187.69$214.14
Assumptions, exclusions and billing units

This is a synthetic scenario, not a customer invoice. The CPU default is rounded from local consumed CPU measurements, roughly 0.12 to 0.15 seconds per transform on this input mix. Deployed CPU usage and idle overhead may differ. Wider images, AVIF sources and different qualities change compute cost.

Every request is conservatively modeled as reaching the Durable Object and R2, with one extra R2 read per transform and one write per fresh derivative. Workers bill both the gateway and the cached loopback invocation. The assumed 7 ms CPU per external request is combined across both entrypoints. R2 holds a rolling 30 days of unique outputs in Standard storage, with the required lifecycle rule. Container output transfer uses North American rates. Durable Objects are assumed active for the entire container awake period at 128 MB each. Rounded R2 and Durable Object billing units and available account allowances are applied.

Awake time cannot fall below the CPU time needed to process the modeled volume. The calculator raises it to that theoretical minimum when needed. That floor assumes full CPU utilization and is not a throughput promise. Uniform fresh traffic can keep every shard awake because of the two-minute idle timeout.

Vercel writes are fresh transforms multiplied by output size rounded up to 8 KiB units. Enter billed cache read units from your dashboard; they are not image request counts. Recently cached regional reads can avoid that meter. Prices use on-demand regional rates before plan credits.

Excluded: Vercel subscription, credits, CDN requests and data transfer; source-host transfer charges on either provider; excess logs/traces, Durable Object storage and container idle/startup CPU overhead. Taxes and support time are excluded. Actual cache reuse, unique retained objects, geography and shared account usage change your bill. The chart does not promise savings or include every platform charge.

Workers Cache also bills internal fetch invocations. See cache invocation pricing. Rates checked October 5, 2026. Provider prices can change.

04 / METHODOLOGY

Know what was measured.

Inputs and verification

50 public PNG sources, 1.1 to 4.3 megapixels, averaging about 3.2 MB. Width 640, quality 98, WebP effort 4, one libvips thread per image. Source query parameters created fresh cache keys without changing bytes. Successful outputs were decoded and dimensions checked.

Live delivery

Open-loop arrivals at 20/s for 60 seconds, separate simultaneous bursts, then cache repeats. Latency runs from request start to the complete response body. Sequential provider runs from one client; no multi-region SLO or long-term reliability claim.

Local capacity

Actual Worker routing and coordinator modules upload to Sharp over HTTP in Docker with enforced CPU and memory limits and swap disabled. Source bytes are replayed locally and R2 writes are simulated with 300 ms latency. Timings include output validation. Each smaller workload ran once for 30 seconds.

Bounded work

Eight active uploads and at most 128 distinct jobs per shard, with a 20 s deadline. Native cancellation retains its slot until the operation settles. Failed configurations stay visible; latency for only successful requests does not demonstrate capacity when failures occur.