<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Publication:
A computationally efficient piecewise linear training algorithm for neural networks utilizing continuous special ordered sets

Loading...
Thumbnail Image

School / College / Institute

Item type:Organizational Unit,
Item type:Organizational Unit,

Program

Organization Authors

Co-Authors

Date

Language

eng

Embargo Status

N/A

Journal Title

Journal ISSN

Volume Title

Alternative Title

Abstract

Artificial neural networks are commonly employed for data-driven modelling of complex nonlinear processes; however, their training may be hindered by the nonlinearity of activation functions and the dependence on local solvers. Obtaining an efficient global solution for neural network training continuous to be an unresolved challenge. A common strategy involves approximating activation functions using piecewise linear formulations to convexify the problem, however this often results in high computational costs due to the addition of auxiliary binary variables. This study proposes integrating piecewise linear formulations for neural network activation functions with a tailored branching algorithm that explores efficient linear programming relaxations to effectively explore the solution space. In the proposed framework, a subset of network parameters is obtained via regular, gradient-descent based training and kept fixed during the training process, while the remaining parameters are determined through the proposed formulation. Unlike conventional mixed-integer programming approaches, where the reliance on binary variables increases the computational complexity, the applied continuous special ordered set method achieves polynomial growth in computation, thereby ensuring better scalability for larger problem instances. The proposed method achieves minimal training error while significantly reducing CPU time. Experiments on datasets of equal dimensionality confirm that the efficiency of the algorithm is not dataset-specific, demonstrating consistent CPU time trends across multiple datasets. These findings highlight the generalizability of the suggested method and its potential to enhance artificial neural network training efficiency across various applications.

Source

Publisher

Elsevier

Citation

item.page.haspartof

Source

Chemical Engineering Research and Design

item.page.ispartofseries

item.page.edition

DOI

10.1016/j.cherd.2026.03.047

item.page.datauri

item.page.link

Rights

N/A

Copyrights Note

Rights and licensing

Endorsement

Review

Supplemented By

Referenced By

Related Patent

Related Goal

Google Scholar
Scholar'da Ara ↗
0
Görüntülenme
0
İndirme
Altmetric
Dimensions
PlumX Metrikleri
BIP! Indicators