Aldi Muhamad Fitrah, Nanang Husin
While Multimodal Artificial Intelligence (AI) systems can assess basic caption image consistency, they lack the visual understanding required to detect complex forgeries. This research used a pipeline that combines a ViT-based anomaly triage and a visually grounded dialogue using an LLaVA family VLM with SAM-based masking based on photographer-guided experience. The proposed model integrates the perceptual expertise of a photographer with the analytical power of modern Vision Language Models (VLM) and Segmentation Tools to facilitate an interactive, evidence-based analysis. The data taken from CASIA v2.0 and COVERAGE evaluation shows novice accuracy improving from 48% (unaided) and 55% (basic multimodal) to 85% with the proposed framework. This photographer-guided approach improves the ability of non-experts to detect and articulate image manipulations compared to both unaided analysis and basic automated multimodal methods. © 2025 IEEE.
Concentration in Journalism Pasundan University, Department of Photography, Indonesia; State University of Surabaya Surabaya, Department of Digital Business, Indonesia