Optimizing Deep Learning for U.S. Technology Stock Price Forecasting Through Data Preprocessing and Hyperparameter Tuning

Authors

  • Fauzan Ghazi Faculty of Artificial Intelligence, Universiti Teknologi Malaysia, Malaysia https://orcid.org/0009-0001-2991-7352
  • Syahid Anuar Faculty of Artificial Intelligence, Universiti Teknologi Malaysia, Malaysia https://orcid.org/0000-0001-7869-0257
  • Saharudin Ismail Faculty of Artificial Intelligence, Universiti Teknologi Malaysia, Malaysia
  • Mariam Mazlan Azman Hashim International Business School, Universiti Teknologi Malaysia, Malaysia

DOI:

https://doi.org/10.47852/bonviewJDSIS620210864

Keywords:

deep learning, stock price forecasting, LSTM, CNN, hyperparameter tuning

Abstract

This study examines how data preprocessing and hyperparameter tuning affect Long Short-Term Memory (LSTM) and Convolutional Neural Network (CNN) models for forecasting selected U.S. technology stock prices. Daily Open, High, Low, Close, and Volume (OHLCV) data were collected for seven major technology companies over ticker-specific periods between 2000 and 2025. The models were evaluated using baseline, feature-augmented, Grid Search, and Optuna-tuned configurations. The results show that structured preprocessing and tuning influenced performance, although the effects differed across architectures and stocks. Adding technical indicators did not consistently improve forecasting accuracy and sometimes increased error, particularly in feature-augmented CNN configurations. Optuna achieved lower observed errors than Grid Search in several cases, while Grid Search remained superior for some tickers. LSTM configurations generally achieved stronger aggregate variance explanation, while CNN with raw inputs and Grid Search achieved the lowest aggregate Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE). CNN produced higher exploratory Sharpe values in several configurations, but these values were calculated without a complete trading-cost and execution framework. The study contributes an empirical comparison of preprocessing and tuning choices within LSTM and CNN forecasting pipelines. The findings are limited to the selected stocks, fixed validation and test periods, and evaluated neural configurations and do not establish statistical superiority or realizable trading performance.

 

Received: 13 June 2026 | Revised: 27 July 2026 | Accepted: 19 August 2026

 

Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.

 

Data Availability Statement

The data that support the findings of this study are openly available on Google Drive at https://drive.google.com/drive/folders/1F71VKe468V0iVB5cvm5kt5sS25DRKoXk?usp=sh aring. The original stock-price data were obtained from the Yahoo Finance historical-data pages for Apple (AAPL), https://finance.yahoo.com/quote/AAPL/history/; Amazon (AMZN), https://finance.yahoo.com/quote/AMZN/history/; Microsoft (MSFT), https://finance.yahoo.com/quote/MSFT/history/; NVIDIA (NVDA), https://finance.yahoo.com/quote/NVDA/history/; MetaPlatforms (META), https://finance.yahoo.com/quote/META/history/; Tesla (TSLA), https://finance.yahoo.com/quote/TSLA/history/; and Alphabet (GOOGL), https://finance.yahoo.com/quote/GOOGL/history/. The contextual macroeconomic series were obtained from Federal Reserve Economic Data (FRED), including the CBOE Volatility Index (VIXCLS), https://fred.stlouisfed.org/series/VIXCLS; Consumer Price Index for All Urban Consumers (CPIAUCSL), https://fred.stlouisfed.org/series/CPIAUCSL; Federal Funds Effective Rate (FEDFUNDS), https://fred.stlouisfed.org/series/FEDFUNDS; Real Gross Domestic Product (GDPC1), https://fred.stlouisfed.org/series/GDPC1; and Unemployment Rate (UNRATE), https://fred.stlouisfed.org/series/UNRATE. The repository contains the raw OHLCV data, feature-engineered datasets, chronological data splits, and collected macroeconomic series. The macroeconomic series were retained as contextual data but were not included as direct inputs in the final LSTM and CNN experiments.

 

Author Contribution Statement

Fauzan Ghazi: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization, Project administration. Syahid Anuar: Methodology, Validation, Writing – review & editing, Supervision, Project administration, Funding acquisition. Saharudin Ismail: Validation, Writing – review & editing, Supervision, Funding acquisition. Mariam Mazlan: Writing – review & editing.

Downloads

Published

2026-09-11

Issue

Section

Research Articles

How to Cite

Ghazi, F., Anuar, S., Ismail, S., & Mazlan, M. (2026). Optimizing Deep Learning for U.S. Technology Stock Price Forecasting Through Data Preprocessing and Hyperparameter Tuning. Journal of Data Science and Intelligent Systems. https://doi.org/10.47852/bonviewJDSIS620210864