3 results
Ranked from CortexLab’s local platform index and published content.
01
model
→02Contextual Bandit
Chooses among actions using context while learning from immediate reward.
model
→03Q-Learning
An off-policy temporal-difference method that learns action values from reward transitions.
model
→Thompson Sampling
Samples action quality from posterior beliefs to balance exploration and exploitation.