Model Library/Depth Anything 3 (Mono Large)

Depth Anything 3 (Mono Large)

Depth Estimation
ComfyUI

ByteDance's visual-geometry model for monocular and video-consistent depth estimation, recovering spatially consistent geometry from any visual input

On-Demand Dedicated 1xRTX 4090

Details

Modalities

image

Recommended Hardware

1xRTX 4090

Estimated Price

Loading...

Provider

ByteDance Seed

Family

Depth Anything

License

Apache 2.0

Depth Anything 3: Visual Geometry from Any Views

Depth Anything 3 (DA3) is the visual geometry model series from the ByteDance Seed team, released in November 2025 as the successor to Depth Anything V2. Where V2 predicts a disparity map from a single image, DA3 predicts depth directly and recovers spatially consistent geometry from arbitrary visual inputs — single images, multi-view sets, or video — with or without known camera poses. This template runs the DA3MONO-LARGE checkpoint, the monocular-depth-optimized member of the family, using the safetensors conversion published at Comfy-Org/Depth-Anything-3.

Key Ideas

The DA3 paper is built on two findings. First, a single plain transformer — a vanilla DINO encoder with no architectural specialization — is a sufficient backbone for visual geometry. Second, a unified depth-ray representation lets one model family cover monocular depth, multi-view depth, camera pose estimation, and 3D reconstruction without complex multi-task learning. The models are trained on public academic datasets only, with a teacher-student pipeline in the tradition of the earlier Depth Anything releases.

Key Features

  • Direct depth prediction rather than V2-style disparity, giving superior geometric accuracy
  • Video-consistent depth: frames processed with cross-view attention stay temporally stable instead of flickering
  • Sky segmentation output alongside the depth map
  • Significantly outperforms Depth Anything 2 on monocular depth estimation, and the wider DA3 family surpasses VGGT on multi-view depth and camera pose benchmarks
  • Compact single-file checkpoint at 0.35B parameters

The Model Family

The series spans Small, Base, Large, and Giant any-view geometry models plus specialized variants: DA3MONO-LARGE for monocular relative depth, DA3METRIC-LARGE for metric depth, and a nested Giant-Large teacher model. The Mono-Large checkpoint served here is the variant ComfyUI's official Depth Anything 3 workflows are built around, tuned for the highest-quality single-image and per-frame depth.

Use Cases

  • Depth maps for 3D parallax, relighting, depth-of-field, and compositing effects
  • Temporally consistent depth for video editing and VFX pipelines
  • Depth conditioning for image and video generation workflows
  • Robotics, drone, and navigation prototyping from a monocular camera signal
  • Preprocessing for point clouds, novel view synthesis, and 3D reconstruction

On This Template

The template runs Depth Anything 3 inside ComfyUI using the native DA3 nodes, with the two official workflows preloaded. The image workflow takes any picture and produces a normalized depth map; the video workflow runs consistent depth over every frame of an uploaded clip and renders the result back to video. Outputs can be wired into other ComfyUI nodes or exported for use in external tools.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.