An AI-Driven Multi-Objective Framework for Interpretable Near-Infrared Spectral Preprocessing

Authors

DOI:

https://doi.org/10.47852/bonviewAIA620210969

Keywords:

automated preprocessing, multi-objective optimization, near-infrared spectroscopy, Pareto optimization, interpretable machine learning

Abstract

Preprocessing has a large, dataset-dependent influence on near-infrared (NIR) calibration, yet analysts often choose transformation sequences by trial and error. This makes the choice difficult to reproduce and may add steps that yield little practical benefit. We therefore cast preprocessing selection as a constrained bi-objective search in which calibration error and the number of active operations are minimized together. Candidate pipelines combined standard normal variate correction, multiplicative scatter correction, detrending, and Savitzky–Golay smoothing or derivatives, subject to chemometric ordering constraints. Random Search provided a non-guided control; Non-dominated Sorting Genetic Algorithm II (NSGA-II) and Multi-objective Tree-structured Parzen Estimator (MOTPE) were evaluated with an identical budget of 200 trials on six public NIR and visible-NIR datasets. All search decisions used calibration data, after which one Pareto-efficient pipeline was refitted and tested once on the held-out subset. NSGA-II gave the better mean validation-root mean square error of prediction (RMSEP) rank (1.33), whereas MOTPE gave the better mean normalized-hypervolume rank (1.25). Compared with the strongest manual reference for each dataset, the optimized pipelines lowered validation RMSEP in five cases by between 3.76% and 16.88%. Wheat was the exception: manual Standard normal variate (SNV) generalized better. Across 10 seeds, held-out RMSEP varied little for most datasets even when selected pipelines differed. Nonlinear calibration was chiefly useful for Mango. Adding wavelength selection reduced Wheat RMSEP to 0.5829 and removed 28.7% of its spectral variables. These findings position multi-objective search as an auditable aid for choosing concise preprocessing pipelines, not as a substitute for independent validation or chemometric judgment.

 

Received: 20 June 2026 | Revised: 17 August 2026 | Accepted: 4 September 2026

 

Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.

 

Data Availability Statement

The data that support the findings of this study are openly available in the Eigenvector Research Inc. Standardization Benchmark Archives (Corn dataset) at http://www.eigenvector.com/data/Corn, in the Eigenvector Research Data Repository (CGL_NIR dataset) at https://eigenvector.com/resources/data-sets/, in Zenodo (Wheat dataset) at https://doi.org/10.5281/zenodo.16759587, in Mendeley Data (Mango dataset) at https://doi.org/10.17632/46htwnp833.3, in Mendeley Data (Soil Organic Matter dataset) at https://doi.org/10.17632/yt78nwnhbd.1, and in Data in Brief (Milk dataset) at https://doi.org/10.1016/j.dib.2023.109767.

 

Author Contribution Statement

Ali Khumaidi: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing – original draft, Visualization, Supervision, Project administration. Ridwan Raafi’udin: Software, Validation, Formal analysis, Data curation, Writing – review & editing.


Downloads

Published

2026-09-17

Issue

Section

Research Article

How to Cite

Khumaidi, A., & Raafi’udin, R. (2026). An AI-Driven Multi-Objective Framework for Interpretable Near-Infrared Spectral Preprocessing. Artificial Intelligence and Applications. https://doi.org/10.47852/bonviewAIA620210969