Manipulator Control Using DRL: Sim-to-Real Validation on a 3-DoF Arm

Authors

  • Shivkumar Sankaralingam Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham-Bengaluru, India https://orcid.org/0009-0000-0028-5163
  • Nippun Kumaar Arulmani Angamuthu Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham-Bengaluru, India and Department of Computer Science and Engineering (AI &ML), Dayananda Sagar University-Bengaluru, India https://orcid.org/0000-0001-5576-0400
  • Suja Palaniswamy Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham-Bengaluru, India https://orcid.org/0000-0001-8252-5828

DOI:

https://doi.org/10.47852/bonviewJCCE62027656

Keywords:

deep reinforcement learning, DDPG, TD3, manipulator control, simulation-to-real transfer

Abstract

Deep reinforcement learning is emerging as a powerful alternative to traditional inverse kinematics for controlling robotic manipulators. By learning optimal actions through interaction with the environment, it enables adaptable and precise control in complex continuous workspaces, making it suitable for dynamic manipulator operations. This paper proposes an actor–critic deep reinforcement learning framework for manipulator control in a continuous workspace. Conventional inverse kinematics solutions can become computationally complex for high-degree-of-freedom manipulators or complex kinematic structures. Accurately generating joint angles by deriving configuration-specific equations therefore remains a persistent challenge. This study employs deep deterministic policy gradient and twin delayed deep deterministic policy gradient algorithms to directly predict joint angles for target-reaching tasks, eliminating reliance on conventional inverse kinematics formulations. A key contribution is a simple yet effective linear reward function that supports stable convergence in continuous action spaces. The proposed framework is implemented using ROS2 and Gazebo simulation and validated on a custom-built physical manipulator. The work addresses complex equation-based control through reinforcement learning and demonstrates proof of concept on a 3-degree-of-freedom manipulator. Experimental results show high task success rates of 95.77% for the deep deterministic policy gradient algorithm and 94.71% for the twin delayed deep deterministic policy gradient algorithm, with accurate end-effector positioning within 0.01 m. These results confirm the effectiveness of the proposed framework for accurate and reliable manipulator control with reduced model complexity and improved real-world applicability.



Received: 13 September 2025 | Revised: 14 April 2026 | Accepted: 10 June 2026



Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.



Data Availability Statement

Data are available from the corresponding author upon reasonable request.



Author Contribution Statement

Shivkumar Sankaralingam: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing, Visualization. NippunKumaar Arulmani Angamuthu: Conceptualization, Methodology, Formal analysis, Investigation, Writing – review & editing, Visualization, Supervision. Suja Palaniswamy: Writing – review & editing, Visualization, Project administration, Funding acquisition.

Downloads

Published

2026-07-17

Issue

Section

Research Articles

How to Cite

Sankaralingam, S., Arulmani Angamuthu, N. K., & Palaniswamy, S. (2026). Manipulator Control Using DRL: Sim-to-Real Validation on a 3-DoF Arm. Journal of Computational and Cognitive Engineering. https://doi.org/10.47852/bonviewJCCE62027656