ECCV 2026

Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

We present RL-AWB, the first framework that integrates reinforcement learning into automatic white balance for nighttime color constancy.
1MediaTek Inc. 2National Taiwan University 3National Yang Ming Chiao Tung University

*Indicates equal contribution

White-balanced output produced by RL-AWB
Low-light input image before white-balance correction
Input RL-AWB

Drag to compare the low-light input with the RL-AWB correction.

Overview

Nighttime white balance as sequential decision-making

Comparison of traditional AWB tuning and the proposed RL-AWB framework
Figure

Nighttime color constancy remains a challenging problem in computational photography due to low-light noise and complex illumination conditions that limit cross-sensor generalization. We present RL-AWB, a novel framework combining statistical methods with deep reinforcement learning for nighttime white balance. Our method begins with a statistical algorithm tailored for nighttime scenes, integrating salient gray pixel detection with novel illumination estimation. Building on this foundation, we develop the first deep reinforcement learning approach for color constancy that leverages the statistical algorithm as its core, mimicking professional AWB tuning experts by dynamically optimizing parameters for each image without knowing its ground-truth illumination. To facilitate cross-sensor evaluation, we introduce the first multi-sensor nighttime color constancy dataset. Experiment results demonstrate that our method achieve superior generalization capability across low-light and well-illuminated images.

What this work contributes

  • We develop SGP-LRD (Salient Gray Pixels with Local Reflectance Differences), a nighttime-specific color constancy algorithm that achieves state-of-the-art illumination estimation on public nighttime benchmarks.
  • We design the RL-AWB framework with Soft Actor-Critic (SAC) training and two-stage curriculum learning, enabling adaptive per-image parameter optimization with exceptional data efficiency.
  • We contribute LEVI (Low-light Evening Vision Illumination), the first multi-camera nighttime dataset comprising 700 images from two sensors, enabling rigorous cross-sensor color constancy evaluation.
  • Extensive experiments demonstrate superior cross-sensor generalization over state-of-the-art with only 5 training images per dataset.

Method

RL-AWB

A Soft Actor-Critic agent observes illumination- and parameter-related features, then iteratively adapts SGP-LRD for each image.

Overview of the proposed RL-AWB framework
Framework overview. SGP-LRD supplies a nighttime-specific estimator; the policy and value networks learn image-adaptive parameter adjustments through a two-stage curriculum.

Adaptive process

Correction improves step by step

The complete trajectory shows the raw input, SGP-LRD initialization, three RL updates, and the corresponding illuminant path.

Images are gamma-corrected for visualization.

Raw input
Raw input for RL-AWB adaptive-process scene 1
SGP-LRD
SGP-LRD for RL-AWB adaptive-process scene 1
RL Step 1
RL Step 1 for RL-AWB adaptive-process scene 1
RL Step 2
RL Step 2 for RL-AWB adaptive-process scene 1
RL Step 3
RL Step 3 for RL-AWB adaptive-process scene 1
Tuning process
Tuning process for RL-AWB adaptive-process scene 1

Dataset

The LEVI dataset

700 nighttime RAW images from two camera systems, with manually annotated ColorChecker masks for cross-sensor evaluation.

Images are gamma-corrected for visualization.

iPhone 16 Pro 370 images 12-bit linear RAW
Sony ILCE-6400 330 images 14-bit linear RAW
Capture range ISO 500-16,000 Low-light nighttime scenes
Illuminant distributions of NCC and LEVI
Broader illuminant coverage. LEVI complements NCC across challenging nighttime lighting conditions.
Normalized RAW mean luminance distributions of NCC and LEVI
More low-luminance scenes. The distribution emphasizes the operating regime targeted by RL-AWB.

Qualitative

Compare corrections, method by method

Select a method to compare its corrections across all four scenes. Reported numbers are angular errors in degrees.

Images are gamma-corrected for visualization.

Scene 01

Raw input for qualitative scene 1
Raw input
RL-AWB output for qualitative scene 1
RL-AWB

Scene 02

Raw input for qualitative scene 2
Raw input
RL-AWB output for qualitative scene 2
RL-AWB

Scene 03

Raw input for qualitative scene 3
Raw input
RL-AWB output for qualitative scene 3
RL-AWB

Scene 04

Raw input for qualitative scene 4
Raw input
RL-AWB output for qualitative scene 4
RL-AWB

Quantitative

Results

Evaluation of the statistical-based methods on both NCC and LEVI datasets. Angular error in degrees.

Cells highlighted as 1st / 2nd / 3rd per column.

Method NCC LEVI
Med.MeanTri.B-25W-25 Med.MeanTri.B-25W-25
GE-1st 4.145.174.351.2510.87 3.944.313.971.827.45
GE-2nd 3.584.643.781.119.93 4.174.494.191.807.76
MSGP 2.483.522.700.808.02 3.123.343.141.545.52
GI 3.134.523.400.9110.60 3.103.423.151.495.91
BCC 3.063.813.231.057.78 4.234.534.282.527.06
RGP 2.223.332.440.687.81 3.213.563.291.636.12
SGP-LRD (Ours) 2.123.112.290.687.22 3.083.253.071.405.46

Cross-dataset evaluation of the learning-based methods between NCC and LEVI datasets. All learning-based baselines are implemented using three-fold cross-validation protocols and trained on the complete dataset.

Cells highlighted as 1st / 2nd / 3rd per column.

Train → Test NCC → LEVI LEVI → NCC
Method Med.MeanTri.B-25W-25 Med.MeanTri.B-25W-25
FC4 11.812.311.96.3619.2 13.214.413.86.1724.1
FFCC 4.447.125.171.9616.7 7.548.517.604.3614.8
C4 3.093.473.181.186.37 5.856.736.042.8612.2
C5 9.1211.79.853.7623.4 4.475.464.681.7010.9
PCC 11.112.511.77.3619.7 8.9610.49.356.2416.7
GCC 28.143.040.512.290.0 9.7741.128.12.2690.0
ePCC 9.6410.29.727.0414.6 7.898.938.175.4614.0
RL-AWB (Ours) 3.033.243.041.455.36 1.993.122.250.677.39

Evaluation results on Gehler-Shi dataset trained on NCC dataset. FC4, C4, C5, PCC, GCC, ePCC and the proposed RL-AWB are trained on NCC dataset and evaluated on Gehler-Shi dataset. Compared with our SGP-LRD, the proposed RL-AWB framework achieves a reduction of 5.9% in the median angular error and 9.8% in the best-25% angular error, showing that RL-AWB generalizes well across low-light and well-illuminated images.

Cells highlighted as 1st / 2nd / 3rd per column.

Method Med.MeanTri.B-25W-25
FC413.815.814.66.3818.8
C45.626.525.842.4312.0
C5 3.343.973.431.327.80
PCC4.978.506.071.6420.9
GCC8.4420.29.683.0658.9
ePCC3.935.164.172.0010.5
SGP-LRD (Ours) 2.383.642.640.518.89
RL-AWB (Ours) 2.243.502.510.468.67
**Qualitative comparison of cross-dataset performance.** Angular error in degrees. Note that images shown are gamma-corrected for visualization.

Qualitative comparison of cross-dataset performance. Angular error in degrees. Note that images shown are gamma-corrected for visualization.

**Qualitative comparison of cross-dataset performance.** Angular error in degrees. Note that images shown are gamma-corrected for visualization.

Qualitative comparison of cross-dataset performance. Angular error in degrees. Note that images shown are gamma-corrected for visualization.

Citation

BibTeX

@article{lee2026rl,
  title={RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes},
  author={Lee, Yuan-Kang and Chen, Kuan-Lin and Chang, Chia-Che and Liu, Yu-Lun},
  journal={arXiv preprint arXiv:2601.05249},
  year={2026}
}