Oscillation of Episode Q0 during DDPG training

3 views (last 30 days)

Show older comments

Heesu Kim on 6 Apr 2021

0
Link

Direct link to this question

https://nl.mathworks.com/matlabcentral/answers/794607-oscillation-of-episode-q0-during-ddpg-training

Commented: Heesu Kim on 6 Apr 2021

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

1 Comment
Show -1 older commentsHide -1 older comments

Heesu Kim on 6 Apr 2021

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

Answers (0)

Products

Reinforcement Learning Toolbox

Release

R2021a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Oscillation of Episode Q0 during DDPG training

1 Comment
Show -1 older commentsHide -1 older comments

Answers (0)

See Also

Categories

Tags

Products

Release

Community Treasure Hunt

Oscillation of Episode Q0 during DDPG training

1 Comment Show -1 older commentsHide -1 older comments

Answers (0)

See Also

Categories

Tags

Products

Release

Community Treasure Hunt

1 Comment
Show -1 older commentsHide -1 older comments