akhaliq HF Staff Claude commited on
Commit
5d4d3b8
·
1 Parent(s): d8a630b

Tighten ZeroGPU reservation with a dynamic duration

Browse files

generate_image previously used @spaces.GPU with the default 60s reserve,
but Z-Image-Turbo at 9 steps finishes in ~3-5s. ZeroGPU reserves GPU time
upfront based on the declared duration, so over-reserving 60s gives this
Space worse queue priority and holds GPUs longer than needed — under high
traffic that pile-up surfaces as 429 'Too many requests' to callers.

Add get_duration(num_inference_steps) estimating steps*0.75s + 3s overhead
(slightly over-estimated so runs never get killed mid-flight) and apply it
via @spaces.GPU(duration=get_duration). Per HF ZeroGPU docs, a shorter
accurate duration raises queue priority and frees GPUs faster, improving
throughput for both this Space's direct users and the workflow frontends
that call /generate_image via gradio_client.

Co-Authored-By: Claude <noreply@anthropic.com>

Files changed (1) hide show
  1. app.py +17 -1
app.py CHANGED
@@ -18,7 +18,23 @@ pipe.to("cuda")
18
 
19
  print("Pipeline loaded!")
20
 
21
- @spaces.GPU
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
  def generate_image(prompt, height, width, num_inference_steps, seed, randomize_seed, progress=gr.Progress(track_tqdm=True)):
23
  """Generate an image from the given prompt."""
24
  if randomize_seed:
 
18
 
19
  print("Pipeline loaded!")
20
 
21
+ def get_duration(prompt, height, width, num_inference_steps, seed, randomize_seed, progress=gr.Progress(track_tqdm=True)):
22
+ """Estimate GPU runtime (seconds) for ZeroGPU scheduling.
23
+
24
+ Z-Image-Turbo is a distilled model: runtime is dominated by the number of
25
+ DiT forwards (num_inference_steps). On a `large` (half Blackwell) GPU each
26
+ step is ~0.3-0.5s. We estimate steps*0.75s plus ~3s fixed overhead for
27
+ kernel warmup + image decode, intentionally over-estimating so ZeroGPU
28
+ never under-reserves and kills a run. The default 60s reserve is far
29
+ longer than needed; a tighter, accurate duration raises this Space's
30
+ queue priority (per HF ZeroGPU docs) and frees GPUs faster, which under
31
+ high traffic reduces the pile-up that surfaces as 429 "Too many requests".
32
+ """
33
+ steps = int(num_inference_steps) if num_inference_steps else 9
34
+ return max(8, int(steps * 0.75 + 3))
35
+
36
+
37
+ @spaces.GPU(duration=get_duration)
38
  def generate_image(prompt, height, width, num_inference_steps, seed, randomize_seed, progress=gr.Progress(track_tqdm=True)):
39
  """Generate an image from the given prompt."""
40
  if randomize_seed: