MoGe-3 ViT-L โ€” MLX (fp32)

Microsoft's MoGe-3 monocular geometry model converted to MLX safetensors for Apple Silicon. 607 tensors, float32.

Usage

from mlx_vlm import load
from mlx_vlm.extraction import extract

model, processor = load("nativ-community/moge-3-vitl-fp32")
result = extract(model, processor, "photo.jpg")
print(result["depth"].shape, result["normal"].shape, result["intrinsics"])

From the command line:

mlx_vlm.extract --model nativ-community/moge-3-vitl-fp32 \
    --image photo.jpg --output geometry.npz

Named outputs are points, depth, normal, mask and intrinsics. Masked-out pixels stay infinite, so read geometry through mask.

Conversion

python -m mlx_vlm.models.moge3.convert \
    --torch-ckpt model.pt --mlx-path moge-3-vitl-fp32 --dtype float32

model.pt comes from Ruicheng/moge-3-vitl. Conversion needs torch; inference does not.

Verification

On a 640x480 photo this checkpoint returns depth (480, 640), normal (480, 640, 3), a boolean mask and 3x3 intrinsics.

Attribution

Weights derive from Ruicheng/moge-3-vitl, MIT licensed, redistributed here converted to MLX under the same licence.

Downloads last month
18
Safetensors
Model size
0.4B params
Tensor type
F32
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nativ-community/moge-3-vitl-fp32

Finetuned
(2)
this model