The Fellowship is a network of exceptional people from different backgrounds who contribute to open-source machine learning ๐งโโ๏ธ๐ฆธโโ๏ธ๐ฆน๐งโโ๏ธ
VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!
Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.
Scribble-Board-Fast is a sketch-to-image workspace powered by Klein-9B, transforming doodles, brush strokes, stickers, and uploaded images into high-fidelity visuals with 4-step distilled sampling.
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.
ImageShield-MMCF โ Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.
If an agent can build the obvious demo, the obvious demo probably isnโt worth building anymore. For years, turning a research repo into something people could actually try was valuable by itself. That part is becoming automated โ and thatโs a good thing.
Which means the interesting work moves elsewhere: finding the weird use case, the right interaction, the unexpected model combination โ or simply knowing which paper is worth anyoneโs attention.
The demo used to be the product. Now it needs a point of view.
๐จ I've just published Sentence Transformers v6.0, introducing MultiVectorEncoder: ColBERT-style late interaction models are now a fourth model type, for training, inference, and interpretation, alongside the dense, sparse, and reranker models! Details:
Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, charts and tables included, with no OCR step in between.
Any PyLate, Stanford ColBERT, or ColPali checkpoint loads straight into the same familiar API: model.encode_query(), model.encode_document(), and model.similarity() just work, whether the documents are texts or page images.
Does it help? LightOn trained LateOn (multi-vector) and DenseOn (dense) on the same data with the same 149M ModernBERT backbone, and the multi-vector model wins on 9 of the 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean NDCG@10. The price is a bigger index, and the new HierarchicalTokenPooling module halves it at roughly no retrieval cost.
Antoine Chaffin, Raphaรซl Sourty, and I wrote a blog post walking through multi-vector models in practice: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Check it out if you want to get started, or just point your Agent to the URL: https://huggingface.co/blog/multi-vector-encoder
**I benchmarked HF buckets against https access for Common Crawl.**
Took me a while to get round to do this but I benchmarked access to Common Crawl via https vs hf buckets. Both experiments were run at night in Europe. I do not think other hardware problems were impacting the speeds since CPU processing time of the non-download pipeline components were highly similar (within 2% identical) and below only the WarcReader speeds of datatrove are used.
Experiment: selected 5 disjoint samples of 64 files each (randomly from the latest crawl; 20,499 docs/file). Those five batches were then processed by 32 single-core tasks with 4GB/core (five batches to calculate CIs). Paired experiment between using https and hf bucket.
That is a difference of about 4x in streaming speed. You'll see that https is also more stable (smaller CI).
I also ran raw throughput tests to the endpoints to measure rate limiting (64MiB transfer at 8/32/128/256 concurrent readers) and rate limiting seems not an issue for either: at any of those parallel reader numbers, their respective speeds stay about the same.
Note that, given CC scale, this is still a small test. Rate limiting may become more obvious when processing a full crawl. I do not know whether the https endpoint vs HF bucket will shut you out earlier with which limits.
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.๐ค
โฑ๏ธ Built a small Space for Visual Chronometer / Pulse of Motion.
Upload a video and estimate its Physical FPS: the frame rate implied by visual motion, independent of metadata. Useful to inspect โchronometric hallucinationโ in generated videos: clips that look smooth, but move with the wrong physical time scale.
A few weeks ago, @victor opened the door: coding agents can now ship Hugging Face Spaces autonomously.
I pulled on that thread.
As someone who builds and ships Gradio demos regularly, I didnโt just want to reproduce the loop. I wanted to see what happens when that loop is plugged into the whole Hugging Face stack.
The interesting part is not only that an agent can ship a Space.
Itโs what happens when Space generation becomes a first-class Hugging Face workflow.
Wan2.2-I2V-Fast with highly upscaled sequential frame sampling is now available as a Spaces demo, built using Wan2.2-I2V and FLUX.2-Klein. Try the demo using the links below.๐
New blog post! An introduction to a little-known but highly effective model reduction method: ๐ง๐ฟ๐ถ๐บ๐บ๐ถ๐ป๐ดโ๏ธ We show how to reduce model size (we went up to 87.24% reduction) while preserving its performance.
We applied this technique to 16 different model families across several modalities to illustrate that it works on any architecture (as long as the embedding layer is the last one of the model) and on any modality involving text. From these 16 families, we generated over ๐ฑ,๐ฑ๐ฌ๐ฌ ๐บ๐ผ๐ป๐ผ๐น๐ถ๐ป๐ด๐๐ฎ๐น ๐บ๐ผ๐ฑ๐ฒ๐น๐ ๐ถ๐ป ๐ญ๐ฎ๐ฐ ๐ฑ๐ถ๐ณ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ ๐น๐ฎ๐ป๐ด๐๐ฎ๐ด๐ฒ๐ ๐
Key takeaways from our experiments: 1๏ธโฃ Trimming does not require a GPU. Our models were obtained on a CPU. 2๏ธโฃ This method scales up to at least 4B parameters (we did not test beyond that). 3๏ธโฃ Trimmed model is smaller than the original while preserving its performance. If you observe a slight performance drop, just fine-tuned to recover or even surpass the original performance. 4๏ธโฃ For an equivalent compute budget, it is better to trim then fine-tune rather than fine-tuning the original model. Since the model is smaller, you can run more epochs/show more data and get in fine a better model than the original. 5๏ธโฃ Trimming is a competitive alternative to distillation and quantization. E.g. we obtained our alternative to DistilBERT in 9 minutes on CPU vs. 90 hours of GPU for the latter. 6๏ธโฃ Trimming could generate reasoning traces in the language of the trimmed model. This could be an alternative to generating traces in English and then translating them into the desired language.
And many other things (such as how much data are needed, the impact of the database used, the order in which it should be done, etc.) are available in the blogpost!
PiD โ Pixel Diffusion Decoder Image Edit Upscale and Image Generation Upscale, an all-in-one demo, is now live on Spaces! Great improvements in realism-based image generation and editing are powered by FLUX.2-Klein, while image generation is paired with Z-Image, and upscaling is enabled by default!