UAV optical imagery provides the detail needed to map road surfaces, roadside vegetation, buildings and traffic participants in road corridors, but the accuracy of such maps is often estimated with random partitions of spatially and temporally autocorrelated frames. This study quantifies how much a random frame split inflates the accuracy of deep semantic segmentation compared with a sequence-based spatial cross validation that holds out complete UAV flight sequences. U-Net, DeepLabV3+ and SegFormer were trained under identical settings on all 42 labeled sequences (420 frames) of the UAVid benchmark of urban street scenes, together with a pixel-based Random Forest reference. Every frame was predicted once under each five-fold protocol, and the protocols were compared with paired sequence-level tests. The random split gave mIoU values of 72.9–74.0%, whereas sequence-based cross validation gave 68.6–69.9%. The difference of 4.1–4.3 percentage points was present in 41 or 42 of the 42 sequences for every model (Holm-adjusted Wilcoxon p < 10⁻¹⁰), overall accuracy was inflated by 2.1–2.6 points, and the fold-to-fold variability of mIoU was two to three times larger under sequence-based validation. The inflation was largest for static cars (10.4–11.4 IoU points), moderate for road and low vegetation, small for buildings and trees and absent for pedestrians, while the ranking of the models did not change. Sequence-based validation therefore gives a more realistic estimate of the accuracy expected on new flights and is recommended for reporting UAV-based road corridor mapping results.