Skip to content

Add Bria fibo-2 - #14991

Draft
muhammadali-afridi wants to merge 1 commit into
huggingface:mainfrom
muhammadali-afridi:ali/fibo2
Draft

muhammadali-afridi wants to merge 1 commit into
huggingface:mainfrom
muhammadali-afridi:ali/fibo2

Conversation

@muhammadali-afridi

Copy link
Copy Markdown

What does this PR do?

Adds fibo-2, Bria's text-to-image and image-editing model.

  • BriaFibo2Transformer2DModel: an 8.1B flow-matching transformer with a single token stream, in the style of Z-Image. Its text encoder is Qwen3-VL-4B.

    • A Perceiver condenses Qwen3-VL's hidden states into gist tokens, which join the image tokens.
    • Five blocks also read Qwen3-VL through gated cross-attention.
    • Images to edit come in as context_latents. Each is patchified by the shared embedder and gets its own RoPE plane.
  • BriaFibo2Pipeline:

    • text-to-image from structured JSON captions;
    • editing of up to 5 images, with an optional mask.

    Two checkpoints share the pipeline:

    • a distilled turbo checkpoint: 4 steps, no guidance;
    • a full checkpoint: 30 steps, guidance 5.
  • scripts/convert_bria_fibo2_to_diffusers.py converts the original checkpoints.

  • Docs and tests.

This is a draft until the weights are public. The repo ids in the docs are placeholders until then.

Verification

I compared the pipeline against Bria's reference inference code, using the original checkpoints, the same seeds and bf16:

  • A100, SDPA:
    • Text-to-image is pixel-identical, for both the turbo and the full checkpoint.
    • Editing is pixel-identical when the text encoder gets the same images as the reference.
  • H100: the reference code runs FlashAttention-3 there. With pipe.transformer.set_attention_backend("_flash_3_varlen_hub"):
    • text-to-image is pixel-identical for both checkpoints;
    • editing is pixel-identical too (4 turbo cases and 2 full-checkpoint cases).
  • Text-encoder images: by default the pipeline thumbnails each image straight to 768 px for Qwen3-VL, as in training. The reference code upsamples 4脳 first, which changes that input slightly.
  • Batching: in float32, a batch of prompts of different lengths matches the separate single-prompt calls. This includes the varlen attention backends; see the self-review notes.
  • Checks:
    • make style, make quality and make fix-copies pass.
    • check_copies, check_dummies, check_support_list, check_forward_call_docstrings and check_return_annotations pass.
  • New tests: they pass on a Mac (CPU/MPS), except the group offloading and layerwise casting memory tests. Those fail the same way for Z-Image and Flux 2 on that machine, because MPS has no CUDA streams and no float8.

Open before ready for review

  • Hosted diffusers-format weights and the final repo ids.
  • The license header of the new files.
  • mask:
    • The pipeline greys out the region to regenerate, which is the convention of Bria's earlier edit models.
    • We're confirming that fibo-2 follows it.
    • The docs sentence on masks will be updated once that's confirmed.
  • Caption score fields: training dropped aesthetic_score and preference_score from captions.
    • The pipeline strips them from edit prompts.
    • Text-to-image prompts are passed as given.
    • To decide: strip them for text-to-image too?
  • Standard vs modular: this is a standard pipeline. Happy to add a modular one if that's preferred.

Self-review notes

I ran the self-review skill on the diff.

Fixed:

  • The transformer forward docstring had no Returns: section, which check_forward_call_docstrings flags.
  • Varlen backends kept the wrong keys.
    • With prompts of different lengths, the padded gist tokens sat between the real gist tokens and the image tokens.
    • The varlen backends (flash_varlen, _flash_3_varlen_hub, sage_varlen) keep the first mask.sum() keys, so they kept the wrong ones.
    • Padded gist tokens now go last. test_padded_batch_matches_single_prompts covers this.
  • Removed dead code: a PyTorch 2.0 check in the attention processor, and casts that changed nothing.
  • __call__ read the first image's size before check_inputs had validated it.

Not fixed, on purpose:

  • Full-graph compile tests are skipped. RopeEmbedder, copied from Z-Image, builds its frequency table under a device context on first use. Z-Image skips these tests for the same reason.
  • Gist counts break the graph under torch.compile(transformer).
    • With a text mask, the gist count per prompt comes from the text lengths through .tolist().
    • Regional compilation (compile_repeated_blocks) is not affected.

Before submitting

  • Did you use an AI agent (Claude Code, Codex, Cursor, etc.) to help with this PR? If so:
    • Did you read the Coding with AI agents guide?
    • Did you run the self-review skill on the diff?
    • Did you share the final self-review notes in the PR description or a comment?
  • Did you read the contributor guideline?
  • Did you read our philosophy doc? (important for complex PRs)
  • Was this discussed/approved via a GitHub issue or the forum? Please add a link to it if that's the case.
  • Did you make sure to update the documentation with your changes?
  • Did you write any new necessary tests?
  • Are you the author (or part of the team) of the model/pipeline (only applicable for model/pipeline related PRs)?

Who can review?

I'll tag reviewers once the weights are public and the open items above are settled.

馃 Generated with Claude Code

fibo-2 is Bria's 8.1B flow-matching DiT, built on the Z-Image block and
conditioned on Qwen3-VL hidden states two ways: a Perceiver resampler turns
them into gist tokens that share the stream with the image tokens, and five
blocks add gated cross-attention to their own bundle of Qwen3-VL layers. It
uses the Flux 2 VAE.

One pipeline class covers both of the model's jobs and both checkpoints:
text-to-image generation, instruction editing of up to five images with an
optional mask, the turbo checkpoint (4 steps, no guidance) and the guided one.

- BriaFibo2Transformer2DModel, with images to edit passed as context latents
- BriaFibo2Pipeline
- a conversion script from the original checkpoint layout
- model and pipeline tests, and docs

Verified against Bria's reference implementation on an A100 and an H100.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation models tests utils pipelines size/L PR with diff > 200 LOC labels Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation models pipelines size/L PR with diff > 200 LOC tests utils

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant