Main Content

nlhwOptions

R2026b

Option set for nlhw

Description

opt = nlhwOptions creates the default option set for nlhw. Use dot notation to customize the option set, if needed.

example

opt = nlhwOptions(Name,Value) creates an option set with options specified by one or more Name,Value pair arguments. The options that you do not specify retain their default value.

example

Examples

collapse all

Create estimation option set for nlhw to view estimation progress, use the Levenberg-Marquardt search method, and set the maximum iteration steps to 50.

opt = nlhwOptions;
opt.Display = 'on';
opt.SearchMethod = 'lm';    
opt.SearchOptions.MaxIterations = 50;

Load data and estimate the model.

load iddata3
sys = nlhw(z3,[4 2 1],idSigmoidNetwork,idPiecewiseLinear,opt);

Create an options set for nlhw where:

  • Initial conditions are estimated from the estimation data.

  • Subspace Gauss-Newton least squares method is used for estimation.

opt = nlhwOptions('InitialCondition','estimate','SearchMethod','gn');

Name-Value Arguments

collapse all

Specify optional pairs of arguments as Name1=Value1,...,NameN=ValueN, where Name is the argument name and Value is the corresponding value. Name-value arguments must appear after other arguments, but the order of the pairs does not matter.

Before R2021a, use commas to separate each name and value, and enclose Name in quotes.

Example: nlhwOptions('InitialCondition','estimate')

Handling of initial conditions during estimation using nlhw, specified as the comma-separated pair consisting of InitialCondition and one of the following:

  • 'zero' — The initial conditions are set to zero.

  • 'estimate' — The initial conditions are treated as independent estimation parameters.

Estimation progress display setting, specified as the comma-separated pair consisting of 'Display' and one of the following:

  • 'off' — No progress or results information is displayed.

  • 'on' — Information on model structure and estimation results are displayed in a progress-viewer window.

Option to normalize estimation data, specified as true or false. If Normalize is true, then the algorithm uses the method specified in NormalizationOptions to normalize the data.

Because saturation, deadzone, and piecewise-linear nonlinearities are physically meaningful, you must use caution to take normalization into account when specifying initial values. However, even when Normalize is true, the software automatically disables normalization for these nonlinearities when you set the property NormalizationOptions.NormalizationMethod to 'auto'.

Option set for configuring normalization, specified as the options shown in the following table. The first option, NormalizationMethod, determines which method the algorithm uses. The default option is 'auto'. In general, for idnlhw models, a setting of 'auto' is equivalent to a setting of 'center'. However, if your model includes any of the nonlinearity estimators that have physically meaningful parameters—idSaturation, idDeadzone, and idPiecewiseLinear—a setting of 'auto' results in the software disabling normalization for those estimators.

Except for 'medianiqr', each specific method in NormalizationMethod has an associated configuration option, such as CenterMethodType when you specify the 'center' method. For more information about these methods, see the MATLAB® function normalize.

Method or Method OptionValueDescriptionDefault
NormalizationMethod'auto'Set method automatically.

'auto'

(equivalent to either'center')

The 'auto' setting disables input and output normalization for idSaturation, idDeadzone, and idPiecewiseLinear nonlinearities.

'center'Center data to have mean 0.
'zscore'z-score with mean 0 and standard deviation 1.
'norm'2-norm.
'scale'Scale by standard deviation.
'range'Rescale range of data to [min,max].
'medianiqr'Center and scale data to have median 0 and interquartile scale of 1.

CenterMethodType (applies to 'center')

'mean'Center to have mean 0.'mean'
'median'Center to have median 0.

ZScoreType (applies to 'zscore')

'std'Center and scale to have mean 0 and standard deviation 1.'std'
'robust'Center and scale to have median 0 and median absolute deviation 1.

ScaleMethodType (applies to 'scale')

'std'Scale by standard deviation.'std'
'mad'Scale by median absolute deviation.
'iqr'Scale by interquartile range.
'first'Scale by first element of data.

NormValue (applies to 'norm')

Positive real valuep-norm, where p is a positive integer.2

Range (applies to 'range')

2-element row vectorRescale range of data to an interval of the form [a b], where a < b.[0 1]

Weighting of prediction error in multi-output model estimations, specified as the comma-separated pair consisting of 'OutputWeight' and one of the following:

  • 'noise' — Optimal weighting is automatically computed as the inverse of the estimated noise variance. This weighting minimizes det(E'*E), where E is the matrix of prediction errors. This option is not available when using 'lsqnonlin' as a 'SearchMethod'.

  • A positive semidefinite matrix, W, of size equal to the number of outputs. This weighting minimizes trace(E'*E*W/N), where E is the matrix of prediction errors and N is the number of data samples.

Options for regularized estimation of model parameters, specified as the comma-separated pair consisting of 'Regularization' and a structure with fields:

Field NameDescriptionDefault
LambdaBias versus variance tradeoff constant, specified as a nonnegative scalar.0 — Indicates no regularization.
RWeighting matrix, specified as a vector of nonnegative scalars or a square positive semi-definite matrix. The length must be equal to the number of free parameters in the model, np. Use the nparams command to determine the number of model parameters.1 — Indicates a value of eye(np).
Nominal

The nominal value towards which the free parameters are pulled during estimation, specified as one of the following:

  • 'zero' — Pull parameters towards zero.

  • 'model' — Pull parameters towards pre-existing values in the initial model. Use this option only when you have a well-initialized idnlhw model with finite parameter values.

'zero'

To specify field values in Regularization, create a default nlhwOptions set and modify the fields using dot notation. Any fields that you do not modify retain their default values.

opt = nlhwOptions;
opt.Regularization.Lambda = 1.2;
opt.Regularization.R = 0.5*eye(np);

Regularization is a technique for specifying model flexibility constraints, which reduce uncertainty in the estimated parameter values. For more information, see Regularized Estimates of Model Parameters.

Numerical search method used for iterative parameter estimation, specified as the one of the values in the following table.

SearchMethodDescription
'auto'

Automatic method selection

A combination of the line search algorithms, 'gn', 'lm', 'gna', and 'grad', is tried in sequence at each iteration. The first descent direction leading to a reduction in estimation cost is used.

'gn'

Subspace Gauss-Newton least-squares search

Singular values of the Jacobian matrix less than GnPinvConstant*eps*max(size(J))*norm(J) are discarded when computing the search direction. J is the Jacobian matrix. The Hessian matrix is approximated as JTJ. If this direction shows no improvement, the function tries the gradient direction.

'gna'

Adaptive subspace Gauss-Newton search

Eigenvalues less than gamma*max(sv) of the Hessian are ignored, where sv contains the singular values of the Hessian. The Gauss-Newton direction is computed in the remaining subspace. gamma has the initial value InitialGnaTolerance (see Advanced in 'SearchOptions' for more information). This value is increased by the factor LMStep each time the search fails to find a lower value of the criterion in fewer than five bisections. This value is decreased by the factor 2*LMStep each time a search is successful without any bisections.

'lm'

Levenberg-Marquardt least squares search

Each parameter value is -pinv(H+d*I)*grad from the previous value. H is the Hessian, I is the identity matrix, and grad is the gradient. d is a number that is increased until a lower value of the criterion is found.

This algorithm requires Optimization Toolbox™ software.

'grad'

Steepest descent least-squares search

'lsqnonlin'

Trust-region-reflective algorithm of lsqnonlin (Optimization Toolbox)

This algorithm requires Optimization Toolbox software.

'patternsearch'

Solver for nonlinearities without well-defined gradients

You can use the patternsearch (Global Optimization Toolbox) solver to find the minimum of a nonlinear function that does not have a well-defined gradient. This solver requires Global Optimization Toolbox software.

'fmincon'

Constrained nonlinear solvers

You can use the sequential quadratic programming (SQP) and trust-region-reflective algorithms of the fmincon (Optimization Toolbox) solver. If you have Optimization Toolbox software, you can also use the interior-point and active-set algorithms of the fmincon solver. Specify the algorithm in the SearchOptions.Algorithm option. The fmincon algorithms might result in improved estimation results in the following scenarios:

  • Constrained minimization problems when bounds are imposed on the model parameters.

  • Model structures where the loss function is a nonlinear or nonsmooth function of the parameters.

  • Multiple-output model estimation. A determinant loss function is minimized by default for multiple-output model estimation. fmincon algorithms are able to minimize such loss functions directly. The other search methods such as 'lm' and 'gn' minimize the determinant loss function by alternately estimating the noise variance and reducing the loss value for a given noise variance value. Hence, the fmincon algorithms can offer better efficiency and accuracy for multiple-output model estimations.

'adam'

Adaptive moment estimation (Adam)

Adam is a first-order adaptive gradient solver. It supports mini-batch operation when data is segmented into multiple frames or batches. For more information, see Adaptive Moment Estimation (Deep Learning Toolbox).

'sgdm'

Stochastic gradient descent with momentum (SGDM)

SGDM is a first-order momentum-based gradient solver. It supports mini-batch operation when data is segmented into multiple frames or batches. For more information, see Stochastic Gradient Descent with Momentum (Deep Learning Toolbox).

'lbfgs'

Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS)

L-BFGS is a quasi-Newton solver that approximates the inverse Hessian using a limited history of curvature pairs. For more information, see Limited-Memory BFGS (Deep Learning Toolbox).

Option set for the search algorithm, specified as the comma-separated pair consisting of 'SearchOptions' and a search option set with fields that depend on the value of SearchMethod.

SearchOptions Structure When SearchMethod Is Specified as 'gn', 'gna', 'lm', 'grad', or 'auto'

Field NameDescriptionDefault
Tolerance

Minimum percentage difference between the current value of the loss function and its expected improvement after the next iteration, specified as a positive scalar. When the percentage of expected improvement is less than Tolerance, the iterations stop. The estimate of the expected loss-function improvement at the next iteration is based on the Gauss-Newton vector computed for the current parameter value.

1e-5
MaxIterations

Maximum number of iterations during loss-function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as Tolerance.

Setting MaxIterations = 0 returns the result of the start-up procedure.

Use sys.Report.Termination.Iterations to get the actual number of iterations during an estimation, where sys is an idtf model.

20
Advanced

Advanced search settings, specified as a structure with the following fields:

Field NameDescriptionDefault
GnPinvConstant

Jacobian matrix singular value threshold, specified as a positive scalar. Singular values of the Jacobian matrix that are smaller than GnPinvConstant*max(size(J)*norm(J)*eps) are discarded when computing the search direction. Applicable when SearchMethod is 'gn'.

10000
InitialGnaTolerance

Initial value of gamma, specified as a positive scalar. Applicable when SearchMethod is 'gna'.

0.0001
LMStartValue

Starting value of search-direction length d in the Levenberg-Marquardt method, specified as a positive scalar. Applicable when SearchMethod is 'lm'.

0.001
LMStep

Size of the Levenberg-Marquardt step, specified as a positive integer. The next value of the search-direction length d in the Levenberg-Marquardt method is LMStep times the previous one. Applicable when SearchMethod is 'lm'.

2
MaxBisections

Maximum number of bisections used for line search along the search direction, specified as a positive integer.

25
MaxFunctionEvaluations

Maximum number of calls to the model file, specified as a positive integer. Iterations stop if the number of calls to the model file exceeds this value.

Inf
MinParameterChange

Smallest parameter update allowed per iteration, specified as a nonnegative scalar.

0
RelativeImprovement

Relative improvement threshold, specified as a nonnegative scalar. Iterations stop if the relative improvement of the criterion function is less than this value.

0
StepReduction

Step reduction factor, specified as a positive scalar that is greater than 1. The suggested parameter update is reduced by the factor StepReduction after each try. This reduction continues until MaxBisections tries are completed or a lower value of the criterion function is obtained.

StepReduction is not applicable for SearchMethod 'lm' (Levenberg-Marquardt method).

2

SearchOptions Structure When SearchMethod Is Specified as 'lsqnonlin'

Field NameDescriptionDefault
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

The value of FunctionTolerance is the same as that of opt.SearchOptions.Advanced.TolFun.

1e-5
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

The value of StepTolerance is the same as that of opt.SearchOptions.Advanced.TolX.

1e-6
MaxIterations

Maximum number of iterations during loss-function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

The value of MaxIterations is the same as that of opt.SearchOptions.Advanced.MaxIter.

20

SearchOptions Structure When SearchMethod Is Specified as 'patternsearch'

Field NameDescriptionDefault
Algorithm

patternsearch optimization algorithm, specified as one of these values:

  • 'classic'

  • 'nups'

  • 'nups-gps'

  • 'nups-mads'

For algorithm details, see How Pattern Search Polling Works (Global Optimization Toolbox) and Nonuniform Pattern Search (NUPS) Algorithm (Global Optimization Toolbox).

For examples of algorithm effects, see Explore patternsearch Algorithms (Global Optimization Toolbox) and Explore patternsearch Algorithms in Optimize Live Editor Task (Global Optimization Toolbox).

'nups'
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

1e-6
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

1e-6
MaxIterations

Maximum number of iterations during loss function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

'100*numberOfVariables', where numberOfVariables is the number of problem variables
UseParallel

Option to enable or disable parallel processing for improved performance, specified as one of these values:

  • "off" — Run in serial on the MATLAB client.

  • "auto" — Use a parallel pool if one is open or if MATLAB can automatically create one. If a parallel pool is not available, run in serial on the MATLAB client.

  • "on" — Use a parallel pool if one is open or if MATLAB can automatically create one. If a parallel pool is not available, throw an error.

If you do not have a parallel pool open and automatic pool creation is enabled, MATLAB opens a pool using the default cluster profile. To use a parallel pool to run computations in MATLAB, you must have Parallel Computing Toolbox™.

Before R2026b: To run in parallel, set UseParallel to true.

"off"

SearchOptions Structure When SearchMethod Is Specified as 'fmincon'

Field NameDescriptionDefault
Algorithm

fmincon optimization algorithm, specified as one of the following:

  • 'sqp' — Sequential quadratic programming algorithm. The algorithm satisfies bounds at all iterations, and it can recover from NaN or Inf results. It is not a large-scale algorithm. For more information, see Sparsity in Optimization Algorithms (Optimization Toolbox).

  • 'trust-region-reflective' — Subspace trust-region method based on the interior-reflective Newton method. It is a large-scale algorithm.

  • 'interior-point' — Large-scale algorithm that requires Optimization Toolbox software. The algorithm satisfies bounds at all iterations, and it can recover from NaN or Inf results.

  • 'active-set' — Requires Optimization Toolbox software. The algorithm can take large steps, which adds speed. It is not a large-scale algorithm.

For more information about the algorithms, see Constrained Nonlinear Optimization Algorithms (Optimization Toolbox) and Choosing the Algorithm (Optimization Toolbox).

'sqp'
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

1e-6
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

1e-6
MaxIterations

Maximum number of iterations during loss function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

100

SearchOptions Structure When SearchMethod Is Specified as 'adam'

Field NameDescriptionDefault
LearnRate

Learning rate, or the step size, used for training, specified as a positive scalar. If the learning rate is too small, then training can take a long time. If the learning rate is too large, then training can be fast but it might reach a suboptimal result, diverge, or oscillate. The learning rate is denoted by α in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

If you specify LearnRateSchedule as "piecewise", then LearnRate is the learning rate before any scheduled drops.

0.001
GradientDecayFactor

Exponential decay rate of gradient moving average for the Adam solver, specified as a positive scalar less than 1. It controls the smoothing of the exponentially decaying average of past gradients. The gradient decay rate is denoted by β1 in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

If the value of GradientDecayFactor is closer to 1, then the smoothing increases. If the value of GradientDecayFactor is closer to 0, then recent gradients have more impact on the training.

0.9
SquaredGradientDecayFactor

Exponential decay rate of squared gradient moving average for the Adam solver, specified as a positive scalar less than 1. It controls the smoothing of the exponentially decaying average of past squared gradients. The squared gradient decay rate is denoted by β2 in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

Larger values of SquaredGradientDecayFactor adapt more slowly but provide a more stable variance estimate.

0.999
EpsilonSmall constant for numerical stability, specified as a positive scalar. To avoid division by zero when updating network parameters, the solver adds this constant to the denominator. Epsilon is denoted by ϵ in the Adaptive Moment Estimation (Deep Learning Toolbox) section.1e-8
MaxEpochsMaximum number of parameter updates or iterations to use for training, specified as a nonnegative integer. If you specify MaxEpochs as 0, the software disables iterations and only runs initialization or post-processing.200
MaxFunctionEvaluationsMaximum number of objective function evaluations, specified as a positive integer.intmax
LearnRateSchedule

Learning rate schedule type, specified as "none" or "piecewise".

  • "none" — No learning rate schedule. This schedule keeps the learning rate constant and equal to LearnRate.

  • "piecewise" — Piecewise learning rate schedule. This schedule multiplies the learning rate by LearnRateDropFactor every LearnRateDropPeriod number of iterations.

"none"
LearnRateDropFactor

Multiplicative factor for dropping the learning rate, specified as a positive scalar less than or equal to 1. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the factor specified by LearnRateDropFactor every time a certain number of iterations passes. Specify the number of iterations using the LearnRateDropPeriod training option.

0.1
LearnRateDropPeriod

Number of iterations in between learning rate drops, specified as a positive integer. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the drop factor every time the number of iterations specified by LearnRateDropPeriod passes. Specify the drop factor using the LearnRateDropFactor training option.

10
MinCostValue

Target objective value, specified as a nonnegative scalar. If the objective value at the current iteration is less than or equal to MinCostValue, the training stops.

0
ModelSelection

Iteration used to return the model parameters, specified as "best" or "last".

If you specify ModelSelection as "best", the software returns the parameters corresponding to the iteration with the lowest objective value. If you specify ModelSelection as "last", the software returns the parameters corresponding to the final iteration.

"best"
AdvancedStructure used to specify advanced search options consisting of these fields:

WeightDecay — Strength of L2 regularization applied to learnable parameters, specified as a nonnegative scalar.

  • If WeightDecay>0 and UseDecoupledWeightDecay=1, weight decay is applied in a decoupled manner before update computations.

  • If WeightDecay>0 and UseDecoupledWeightDecay=0, classical L2 regularization is applied by augmenting gradients before update computations.

To disable this option, specify WeightDecay as 0.

0

UseDecoupledWeightDecay — Flag to control whether weight decay is applied directly to parameters or is folded into the gradient as L2 penalty, specified as a logical scalar.

  • If UseDecoupledWeightDecay=1, decoupled weight decay is applied directly.

  • If UseDecoupledWeightDecay=0, L2 penalty term is added to the gradient.

Decoupling prevents regularization strength from being implicitly modulated by momentum dynamics.

1

ClipGradNorm — Maximum allowed L2 norm of the flattened gradient vector, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps prevent unstable updates when gradients spike, such as in recurrent models or stiff dynamical systems.

To disable this option, specify ClipGradNorm as 0.

0
NormEpsilon — Small constant used to avoid division by zero and underflow when computing safe norms and normalized quantities, especially when ClipGradNorm is enabled, specified as a positive scalar.1e-12

SearchOptions Structure When SearchMethod Is Specified as 'sgdm'

Field NameDescriptionDefault
LearnRate

Learning rate, or the step size, used for training, specified as a positive scalar. If the learning rate is too small, then training can take a long time. If the learning rate is too large, then training can be fast but it might reach a suboptimal result, diverge, or oscillate. The learning rate is denoted by α in the Stochastic Gradient Descent with Momentum (Deep Learning Toolbox) section.

If you specify LearnRateSchedule as "piecewise", then LearnRate is the learning rate before any scheduled drops.

0.01
Momentum

Momentum coefficient, specified as a positive scalar less than or equal to 1. This coefficient controls the contribution of the previous gradient step to the current iteration. The momentum coefficient is denoted by γ in the Stochastic Gradient Descent with Momentum (Deep Learning Toolbox) section.

If the value of Momentum is closer to 1, then the smoothing increases. If the value of Momentum is closer to 0, then the solver behaves closer to the stochastic gradient descent algorithm.

0.95
MaxEpochsMaximum number of parameter updates or iterations to use for training, specified as a nonnegative integer. If you specify MaxEpochs as 0, the software disables iterations and only runs initialization or post-processing.200
MaxFunctionEvaluationsMaximum number of objective function evaluations, specified as a positive integer.intmax
LearnRateSchedule

Learning rate schedule type, specified as "none" or "piecewise".

  • "none" — No learning rate schedule. This schedule keeps the learning rate constant and equal to LearnRate.

  • "piecewise" — Piecewise learning rate schedule. This schedule multiplies the learning rate by LearnRateDropFactor every LearnRateDropPeriod number of iterations.

"none"
LearnRateDropFactor

Multiplicative factor for dropping the learning rate, specified as a positive scalar less than or equal to 1. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the factor specified by LearnRateDropFactor every time a certain number of iterations passes. Specify the number of iterations using the LearnRateDropPeriod training option.

0.1
LearnRateDropPeriod

Number of iterations in between learning rate drops, specified as a positive integer. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the drop factor every time the number of iterations specified by LearnRateDropPeriod passes. Specify the drop factor using the LearnRateDropFactor training option.

10
MinCostValue

Target objective value, specified as a nonnegative scalar. If the objective value at the current iteration is less than or equal to MinCostValue, the training stops.

0
ModelSelection

Iteration used to return the model parameters, specified as "best" or "last".

If you specify ModelSelection as "best", the software returns the parameters corresponding to the iteration with the lowest objective value. If you specify ModelSelection as "last", the software returns the parameters corresponding to the final iteration.

"best"
AdvancedStructure used to specify advanced search options consisting of these fields:

WeightDecay — Strength of L2 regularization applied to learnable parameters, specified as a nonnegative scalar.

  • If WeightDecay>0 and UseDecoupledWeightDecay=1, weight decay is applied in a decoupled manner before update computations.

  • If WeightDecay>0 and UseDecoupledWeightDecay=0, classical L2 regularization is applied by augmenting gradients before update computations.

To disable this option, specify WeightDecay as 0.

0

UseDecoupledWeightDecay — Flag to control whether weight decay is applied directly to parameters or is folded into the gradient as L2 penalty, specified as a logical scalar.

  • If UseDecoupledWeightDecay=1, decoupled weight decay is applied directly.

  • If UseDecoupledWeightDecay=0, L2 penalty term is added to the gradient.

Decoupling prevents regularization strength from being implicitly modulated by momentum dynamics.

1

ClipGradNorm — Maximum allowed L2 norm of the flattened gradient vector, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps prevent unstable updates when gradients spike, such as in recurrent models or stiff dynamical systems.

To disable this option, specify ClipGradNorm as 0.

0
NormEpsilon — Small constant used to avoid division by zero and underflow when computing safe norms and normalized quantities, especially when ClipGradNorm is enabled, specified as a positive scalar.1e-12

SearchOptions Structure When SearchMethod Is Specified as 'lbfgs'

Field NameDescriptionDefault
MaxIterations

Maximum number of quasi-Newton iterations to use for training, specified as a nonnegative integer. Each iteration forms a search direction using the stored curvature pairs and then performs a line search.

If you specify MaxIterations as 0, the software disables iterations and only runs initialization or post-processing.

200
MaxFunctionEvaluationsMaximum number of objective function evaluations, including evaluations performed by line search, specified as a positive integer.intmax
HistorySize

Number of curvature pairs or state updates to store, specified as a positive integer.

The L-BFGS algorithm uses a history of gradient calculations to approximate the Hessian matrix recursively. Larger values of HistorySize can improve the Hessian approximation but will increase memory usage and cost per iteration. For more information, see the Limited-Memory BFGS (Deep Learning Toolbox) section.

10
GradientTolerance

Stopping tolerance on the relative gradient, specified as a positive scalar.

The software stops training when the relative gradient is less than or equal to GradientTolerance.

1e-6
StepTolerance

Stopping tolerance on the step size, specified as a positive scalar. StepTolerance specifies the minimum allowable change in the parameters between successive iterations.

The software stops training when the step that the algorithm takes is less than or equal to StepTolerance.

1e-12
FunctionTolerance

Stopping tolerance on the improvement in the objective value, specified as a positive scalar. FunctionTolerance specifies the minimum required decrease in the objective function value between iterations.

The software stops training when the objective value improvement is less than or equal to FunctionTolerance.

1e-12
LineSearchMethod

Method to find a suitable step size, specified as one of these values:

  • "strong-wolfe" — Search for a step size that satisfies the strong Wolfe conditions (sufficient decrease and strong curvature). This method maintains a positive definite approximation of the inverse Hessian matrix.

  • "weak-wolfe" — Search for a step size that satisfies the weak Wolfe conditions (sufficient decrease and curvature). This method maintains a positive definite approximation of the inverse Hessian matrix. It can accept longer steps.

  • "armijo" — Search for a learning rate that satisfies sufficient decrease conditions only. This method does not maintain a positive definite approximation of the inverse Hessian matrix. It is often more tolerant of noisy gradients but can accept shorter steps.

"strong-wolfe"
MaxNumLineSearchIterationsMaximum number of line search trials per iteration to determine the step size, specified as a positive integer.40
InitialStepSizeStep size for the starting line search trial, specified as a positive scalar.1.0
AdvancedStructure used to specify advanced search options consisting of these fields:
MinStepSize — Smallest step size permitted by line search, specified as a positive scalar. If the step size for a trial goes below this value, the line search fails and the solver can stop or fall back depending on the implementation.1e-16
MaxStepSize — Largest step size permitted by line search, specified as a positive scalar. This upper bound for the trial step size prevents excessively large moves that can cause numerical overflow or objective evaluation failures.1e+16

GradientClipNorm — Maximum allowed L2 norm of the gradient vector used by the quasi-Newton update and line search, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps improve robustness on problems with occasional gradient spikes or poor scaling.

To disable this option, specify GradientClipNorm as 0.

0
WolfeC1 — Armijo condition (sufficient decrease) constant for Wolfe line search, specified as a positive scalar less than 1. Smaller values of WolfeC1 make sufficient decrease easier to satisfy.1e-4
WolfeC2 — Curvature condition constant for Wolfe line search, specified as positive scalar less than 1. Larger values of WolfeC2 make the curvature condition easier to satisfy whereas smaller values enforce a stronger curvature requirement.0.9
ZoomMaxIterations — Maximum number of iterations allowed in the "zoom" procedure of Wolfe line search, specified as a positive integer.40
BacktrackingFactor — Step size shrink factor during backtracking used to reduce trial step sizes when conditions are not satisfied, specified as a positive scalar less than 1. Values closer to 0 shrink the step size more aggressively while values closer to 1 shrink the step size more conservatively.0.5
CurvatureThreshold — Number to control whether a new curvature pair is accepted into the limited-memory history, specified as a positive scalar. Specifying CurvatureThreshold prevents storing nearly singular or noisy curvature information that can destabilize the inverse-Hessian approximation.1e-10

PowellDamping — Number to control the amount of Powell damping applied when the curvature condition is weak, specified as a nonnegative number less than 1. Damping enforces positive curvature and a positive-definite inverse-Hessian approximation.

To disable this option, specify PowellDamping as 0.

0
UseInitialScaling — Flag to control whether the initial inverse-Hessian is scaled each iteration using curvature information, specified as a logical scalar. This scaling often improves practical performance.1

To specify field values in SearchOptions, create a default nlhwOptions set and modify the fields using dot notation. Any fields that you do not modify retain their default values.

opt = nlhwOptions;
opt.SearchOptions.MaxIterations = 50;
opt.SearchOptions.Advanced.RelImprovement = 0.5;

Additional advanced options, specified as the comma-separated pair consisting of 'Advanced' and a structure with fields:

Field NameDescriptionDefault
ErrorThresholdThreshold for when to adjust the weight of large errors from quadratic to linear, specified as a nonnegative scalar. Errors larger than ErrorThreshold times the estimated standard deviation have a linear weight in the loss function. The standard deviation is estimated robustly as the median of the absolute deviations from the median of the prediction errors, divided by 0.7. If your estimation data contains outliers, try setting ErrorThreshold to 1.6.0 — Leads to a purely quadratic loss function.
MaxSizeMaximum number of elements in a segment when input-output data is split into segments, specified as a positive integer.250000

To specify field values in Advanced, create a default nlhwOptions set and modify the fields using dot notation. Any fields that you do not modify retain their default values.

opt = nlhwOptions;
opt.Advanced.ErrorThreshold = 1.2;

Output Arguments

collapse all

Option set for nlhw, returned as an nlhwOptions option set.

Extended Capabilities

expand all

Version History

Introduced in R2015a

expand all

See Also