Abstract:To address accuracy degradation caused by imaging scale decay and large-angle self-occlusion in mid-to-long-range (1.0~3.0 m) gaze measurement, a dual-view method integrating pose guidance and frequency-domain compensation is proposed. First, to overcome the limitations of near-field datasets, a mid-to-long-range dual-view spatiotemporally aligned dataset that preserves natural eye-head dynamic coupling mechanisms is constructed to provide a reliable benchmark. Second, a head pose-guided feature selection module is designed. By leveraging continuous 6D pose representations to generate spatial attention masks, it uses top-down prior constraints to isolate and suppress spatial noise from perspective distortion and occlusion. Subsequently, to tackle severe scale degradation and topological detail loss, a scale-adaptive high-frequency reconstruction compensation module is proposed. It dynamically infers the physical observation distance via inter-ocular pixel spacing to drive scale embeddings, and employs the 2D discrete cosine transform to bottom-up reconstruct attenuated high-frequency ocular boundary features in a bottom-up manner. Finally, an uncertainty-aware weight learning fusion strategy is introduced. While ensuring spatial geometric consistency, it dynamically penalizes unilaterally degraded views to achieve deep cross-view feature complementation. Experimental results demonstrate that the method reduces the mean angular error to 3.52° and 2.75° on the custom and ETH-XGaze datasets, respectively, exhibiting superior measurement accuracy and robustness.