A 7B multi-modal omni-task model for visual understanding, depth/normal estimation, image generation, editing, detection, and OCR.