cover marc toussaint

Marc Toussaint on planning as inference and graphical models

  • cover play_arrow

    PLAY EPISODE


cover marc toussaint
Season 2013
Season 2013
Description arrow_drop_down

Description

What if planning is not about computing value functions but about performing probabilistic inference? Marc Toussaint shows how recasting optimal control as message passing opens new computational pathways for robotics and decision-making.

Subscribe for more from the Convergent Science Network podcast series.

Marc Toussaint presents a theoretical framework that reformulates planning and optimal control as probabilistic inference in graphical models. Rather than iterating backward through Bellman equations to compute value functions, his approach computes both forward and backward messages whose product yields a posterior distribution over actions. This shift in perspective is not merely notational: it leads to genuinely different approximation algorithms, particularly for complex problems like partially observable Markov decision processes and factored state spaces where traditional value function methods struggle.

The conversation traces the intellectual lineage from Kalman’s duality between control and filtering through Bert Kappen’s work on path integrals to Toussaint’s own generalization that operates over joint state-control processes without restrictive assumptions about dynamics or cost structure. A key theoretical achievement is demonstrating that many existing reinforcement learning algorithms emerge as special cases of this unified formulation, providing both theoretical elegance and inherited empirical validation.

Toussaint derives a model-free reinforcement learning algorithm from this framework where the policy is represented as a Boltzmann distribution. Analysis of its fixed-point properties reveals a surprising result: for non-optimal actions, the Boltzmann energy diverges to negative infinity, making them vanishingly improbable, while for optimal actions, it converges exactly to the optimal value function. The framework handles goal conflicts through the natural machinery of probabilistic inference, where inconsistent evidence simply reduces likelihood and the system finds probabilistic compromises.

The episode also explores Toussaint’s robotics applications, where model-based approaches using stochastic relational rules enable robots to generalize from minimal experience. Active exploration strategies that maximize information gain prove essential in the exponentially large state spaces created by relational representations of multi-object environments, allowing a robot that has observed balls rolling to intelligently seek out non-ball-shaped objects to test next.

Tagged as:

About the author call_made

CSN Podcasts

Both the triumphs of humanity and its most evil deeds have resulted from collaboration. In a time where humanity is required to aspire to the former and minimize the latter, the question arises of how collaboration arises and why it fails. Surprisingly, this phenomenon, so central to who we are, is not well understood. Hence, a collaborative effort is required to understand collaboration in its full biological, psychological, sociological, cultural, and economic complexity and to translate this understanding into operational impact. This series of podcasts is one step toward achieving these complementary goals. The Collaboration Podcast presents interviews with people who are central orchestrators of collaboration in various domains including business, government, science, art, health, sustainability, and the military. The discussions were conducted by Prof. Dr. Paul F.M.J. Verschure and members of the Program Advisory Committee of the Ernst Strungmann Forum on Collaboration (https://www.esforum.de/forums/ESF32_Collaboration.html) during 2021 and had the goal to sketch a map of opportunities, challenges, and obstacles in human collaboration. The forum took place in May 2022, and now we would like to share this series of interviews with a broader audience. The full report of the Forum will be published in 2023 by MIT Press. The podcast was produced by the Convergent Science Network (https://www.convergentsciencenetwork.org/). Context: The stability of social systems depends critically on realizing sustainable methods of “collaboration,” yet how and by which means collaboration is achieved is not clearly understood; neither are the conditions or processes that lead to its breakdown or failure. Collaboration can be understood as cooperation between agents toward mutually constructed goals. Part of the reason for our lack of understanding is that the phenomenon of collaboration is, by nature, a highly multidisciplinary problem, and effective research into its complexities has been difficult to achieve across the broad range of scientific and technical disciplines involved. The need for a fundamental understanding of collaboration, however, has become increasingly important. Not only does humankind demand answers as it attempts to address critical challenges at multiple scales (e.g., climate change, migration, enhanced automation, social and economic inequality), but ever-increasing technological and economic means of interconnecting people and societies are disrupting long-established, familiar patterns of how we interact. Radical technological changes that are ongoing have the potential to reshape collaboration in ways that are currently hard to predict or influence (e.g., by altering configurations in interaction, information creation, and modes of communication). On one hand, such changes could disrupt hitherto stable forms of collaboration by affecting critical communication channels and traditional roles, as can be observed in the rapidly changing patterns in governance, commerce, and social interaction. Conversely, technology could lead to the emergence of novel, successful forms of collaboration that deviate from traditional “hierarchical” architectures. Evidence of this can be seen in areas as diverse as highly automated manufacturing plants, the open science movement, collaborative software repositories, user-centered services, and the sharing of economy-based modes of organization. Without a fundamental understanding of the mechanisms, processes, and boundary conditions of collaboration, it is not possible to evaluate or predict which of these possible scenarios are sustainable or even plausible. The Forum “How Collaboration Arises and Why it Fails” (May 8–13, 2022, Location: Frankfurt am Main, Germany) Chairs: Andreas Roepstorff and Paul Verschure Program Advisory Committee: Jenna Bednar, Julia R. Lupp, Bhavani R. Rao , Andreas Roepstorff, Ferdinand von Siemens, and Paul Verschure

More posts

Timestamp

  • fast_forward00:00:03 - This is the Convergent Science Network podcast. Leading researchers in the domain
  • fast_forward00:00:10 - of neuroscience, brain theory and technology are interviewed by Paul Verschoor and Tony Prescott.
  • fast_forward00:00:26 - This is Paul Verschoor with the Convergent Science Network podcast.
  • fast_forward00:00:30 - And today I'm with Mark Toussaint, who's a speaker in our summer school.
  • fast_forward00:00:35 - And Mark, as a physicist, has been presenting his important work in the domain of machine learning.
  • fast_forward00:00:44 - So Mark, you started your presentation with like a little test for your audience,
  • fast_forward00:00:48 - where you showed them a few equations and you want to check out who in the audience
  • fast_forward00:00:52 - would recognize these equations.
  • fast_forward00:00:54 - So why were these equations so important to you? Well, first just to check what
  • fast_forward00:00:59 - kind of background people have.
  • fast_forward00:01:03 - I'm not so sure if they're so important, but I think they are important to eventually
  • fast_forward00:01:08 - understand what the point is of these formulations that I presented and to sort
  • fast_forward00:01:13 - of understand that the third equation on that slide is something that we can
  • fast_forward00:01:17 - get out of these new formulations. Okay.
  • fast_forward00:01:21 - And actually, I think you looked a bit disappointed because it was not a large
  • fast_forward00:01:25 - group of people who immediately jumped up like yes i know this one so yeah there is something um,
  • fast_forward00:01:32 - i mean reinforcement learning is a very very basic model of
  • fast_forward00:01:35 - how behavior could be organized and how behavior could be
  • fast_forward00:01:37 - organized in a way that is goal directed and um
  • fast_forward00:01:41 - although you know there is many many questions whether this helps you
  • fast_forward00:01:44 - understanding the brain of course there's many questions but still
  • fast_forward00:01:47 - this should not um i don't know free you of knowing these
  • fast_forward00:01:51 - things i think it would be good to teach them yeah absolutely um
  • fast_forward00:01:55 - now what you initially emphasized were
  • fast_forward00:02:00 - your your methods for let's say to understand planning as a form of inference
  • fast_forward00:02:06 - yeah so why do you think that that's a useful way to to think about planning
  • fast_forward00:02:12 - first it's an it's an alternative way to think about planning,
  • fast_forward00:02:17 - which is a good thing in itself.
  • fast_forward00:02:20 - I think, you know, turning over problems and looking at the same problems from
  • fast_forward00:02:24 - different perspectives is important.
  • fast_forward00:02:25 - I think the perspective of looking at planning from the perspective of planning.
  • fast_forward00:02:30 - You know, offers different, you know, options of how to think about a computation
  • fast_forward00:02:37 - that actually realizes planning.
  • fast_forward00:02:39 - So more concretely about why planning is inference, I think it does allow you
  • fast_forward00:02:45 - to be creative to use other kinds
  • fast_forward00:02:47 - of computational algorithms that exploit different types of structures.
  • fast_forward00:02:52 - To be more concrete, if you think about how to do planning on factored representations,
  • fast_forward00:02:57 - one way of thinking, the more classical one, would be to have structured representations
  • fast_forward00:03:02 - of the value function, but it will always be the value function, right?
  • fast_forward00:03:06 - Whereas if you roll out that process and sort of formalize it as the whole model
  • fast_forward00:03:11 - as a graphical model and show that inference in a graphical model can solve
  • fast_forward00:03:15 - planning problems, it really leads to different types of algorithms and different types of,
  • fast_forward00:03:20 - approximations that you could use to solve planning.
  • fast_forward00:03:22 - But would that mean you really have removed the value function from the story,
  • fast_forward00:03:27 - or you have redefined it?
  • fast_forward00:03:31 - That's also very interesting. That's a good question.
  • fast_forward00:03:35 - So I have removed it in the computational algorithm itself so that the algorithm
  • fast_forward00:03:40 - does not have to have explicitly a representation of the value function in its
  • fast_forward00:03:43 - memory or the computer doesn't have to.
  • fast_forward00:03:47 - But in some cases, in particular on normal Markov decision processes,
  • fast_forward00:03:51 - the algorithm will compute things which are one-to-one to the value function.
  • fast_forward00:03:56 - For instance, on normal Markov decision processes, the backward messages are
  • fast_forward00:04:00 - equivalent to the value function.
  • fast_forward00:04:05 - But again, all of this is only true on a normal Markov decision process.
  • fast_forward00:04:10 - If you go to the actually interesting problems where you cannot do exact inference
  • fast_forward00:04:14 - or exact value iteration, like in POMDPs or other kind of factored things.
  • fast_forward00:04:21 - The messages do not anymore correspond to value functions. They really correspond
  • fast_forward00:04:26 - to probabilistic posteriors.
  • fast_forward00:04:29 - And it would be unclear how to do the same things that it could do with inference,
  • fast_forward00:04:33 - how you could do the same things with value functions.
  • fast_forward00:04:36 - Right. So in that sense, also if you go to the Bellman equation view,
  • fast_forward00:04:41 - which would be more dominated by a value function, then for you this is also,
  • fast_forward00:04:46 - let's say, a contrast, right? is okay.
  • fast_forward00:04:48 - That's one way to look at planning. And you want to look at alternative models to solve it.
  • fast_forward00:04:55 - That might be more simple or are they just different formulations of the Bellman view?
  • fast_forward00:05:02 - They are different. I think they're different to the Bellman view.
  • fast_forward00:05:06 - So the Bellman view is really a backward.
  • fast_forward00:05:08 - The Bellman equation is a backward thinking.
  • fast_forward00:05:11 - So if you have already solved a small problem, you can sort of backward iterate
  • fast_forward00:05:14 - and sort of move your horizon backward and solve the larger problem and so forth.
  • fast_forward00:05:22 - So in the Bellman view, it's not very symmetric. It doesn't think about,
  • fast_forward00:05:26 - for instance, forward messages or forward information.
  • fast_forward00:05:29 - The inference view does. It computes both, forward and backward messages.
  • fast_forward00:05:33 - And the result is then a product of both of these being the posterior.
  • fast_forward00:05:37 - So I think it is a different view than the Bellman view. But I could argue that
  • fast_forward00:05:42 - I might, let's say, describe your inference-based approach in terms of a Bellman equation.
  • fast_forward00:05:51 - I can show an equivalence or is there really something missing?
  • fast_forward00:05:57 - I don't think that the computational things that are happening during message
  • fast_forward00:06:03 - passing, you could describe as being a form of Bellman equation.
  • fast_forward00:06:09 - So computationally, the things that are being computed and the kind of messages
  • fast_forward00:06:13 - that are being computed are different.
  • fast_forward00:06:16 - As a result, they will also solve somewhat different problems?
  • fast_forward00:06:20 - Well, they lend to different approximations. For instance, as I said, in a POMD piece.
  • fast_forward00:06:26 - Well, in POMD piece, sometimes you can do exact inference as well.
  • fast_forward00:06:30 - But if you do message passing, you can cope differently with large problems
  • fast_forward00:06:36 - or with factor problems as you
  • fast_forward00:06:38 - could do with trying to approximate value functions and iterate over them.
  • fast_forward00:06:42 - Right. So now you mentioned that that's sort of introductory remarks,
  • fast_forward00:06:47 - but you mentioned that on the one end, you found problem-solving from the perspective
  • fast_forward00:06:52 - of inference theoretically interesting,
  • fast_forward00:06:54 - but also you said, look, it can give us a handle on a more large-scale integration
  • fast_forward00:07:00 - and possibly on some form of multidisciplinary interaction.
  • fast_forward00:07:04 - Right. So what's the perspective there exactly? So the,
  • fast_forward00:07:10 - I think generally, actually, probabilistic modeling and also graphical models,
  • fast_forward00:07:15 - there's two things about them.
  • fast_forward00:07:17 - Some people say they're great tools because they work well, but another thing
  • fast_forward00:07:20 - to say about them would be they're nice because they offer one nice language
  • fast_forward00:07:25 - that can be applied on many different disciplines.
  • fast_forward00:07:28 - So graphical models can be used in linguistics, in robotics,
  • fast_forward00:07:32 - for inference, for whatever.
  • fast_forward00:07:34 - It's a bit of a similar thing that I wanted to state with planning as inference.
  • fast_forward00:07:40 - The first point to make is that it's one language that allows you to talk coherently
  • fast_forward00:07:45 - about both decision-making, doing perceptual inference, sensor processing, while also planning.
  • fast_forward00:07:53 - So not only immediate decision-making, but also planning in one coherent set of concepts.
  • fast_forward00:07:59 - And it's actually a rather small set of concepts, just inference,
  • fast_forward00:08:03 - right? I think this is good for communication between disciplines.
  • fast_forward00:08:07 - That's really one point. And as I said, one example of that is that having these
  • fast_forward00:08:14 - concepts that it can show, that they explain planning and other things,
  • fast_forward00:08:19 - allows to build bridges between disciplines.
  • fast_forward00:08:23 - As I said, there is neuroscientists, and I don't have a known opinion,
  • fast_forward00:08:26 - but there is neuroscientists who like to argue that neurocomputation can be
  • fast_forward00:08:30 - abstracted as functionally doing something as Bayesian inference.
  • fast_forward00:08:35 - Or to say it even weaker, that neurosystems could at least also realize something as Bayesian inference.
  • fast_forward00:08:42 - But that is an interesting statement because if now in other fields we can show
  • fast_forward00:08:46 - that Bayesian inference can solve planning problems, we immediately have a hypothesis
  • fast_forward00:08:50 - of how neurosystems could solve planning problems.
  • fast_forward00:08:53 - But this is also a goal in your own work? I wouldn't really say so.
  • fast_forward00:08:58 - So particularly in the last years, I became more and more of a roboticist, really. And, um.
  • fast_forward00:09:07 - So in a positive sense, yes, I am inspired. And I'm inspired by the questions
  • fast_forward00:09:11 - of how actual systems like the brain can solve problems of goal-directed behavior,
  • fast_forward00:09:18 - manipulation, all of these things.
  • fast_forward00:09:20 - But it would be too much to say that I see myself as a researcher,
  • fast_forward00:09:27 - like a brain researcher.
  • fast_forward00:09:27 - Future right okay so then um let's
  • fast_forward00:09:31 - say the the approach that you have a great confidence in at this point in time
  • fast_forward00:09:36 - is what you call stochastic optimal control which is a problem definition not
  • fast_forward00:09:41 - an approach yeah so but why why
  • fast_forward00:09:44 - do you think this is a an important set of problems then to to deal with.
  • fast_forward00:09:52 - That's again a good question. So let me say what I guess many people would,
  • fast_forward00:09:58 - other people would say, right?
  • fast_forward00:10:01 - Stochastic optimal control is a problem definition similar to Markov decision
  • fast_forward00:10:05 - processes with reward and it's a very general one.
  • fast_forward00:10:09 - People like when things are being formulated general and problems being formulated
  • fast_forward00:10:13 - general and allows to sort of you know, describe all kinds of planning problems in there.
  • fast_forward00:10:20 - And it's well, of course, accepted within the control theory itself.
  • fast_forward00:10:25 - And it's quite a lot of work on within machine learning in terms of Markov decision processes.
  • fast_forward00:10:31 - Let me put it the other way. If I would not have shown and embedded my theory
  • fast_forward00:10:36 - to be important or useful for solving MDPs and optimal control,
  • fast_forward00:10:42 - I could have not convinced the whole control theory community or machine learning
  • fast_forward00:10:46 - community that it's worth anything. Okay.
  • fast_forward00:10:49 - So how did you then solve this problem? Solve which problem?
  • fast_forward00:10:53 - Well, stochastic optimal control.
  • fast_forward00:10:56 - You're saying, look, you try to show that your framework is applicable in that domain.
  • fast_forward00:11:02 - So in that sense, I guess you want to have an impact that sort of also delineates
  • fast_forward00:11:06 - your proposal from others.
  • fast_forward00:11:09 - So what are the unique features of what you propose in that context?
  • fast_forward00:11:14 - Why did it work better or why was it more appreciated than alternatives? Right.
  • fast_forward00:11:19 - Um, I mean, one of the, I mean, achievements, I think, is a theoretical achievement
  • fast_forward00:11:25 - in the sense that we could show that many, many other algorithms,
  • fast_forward00:11:27 - which themselves have, you know, proven also,
  • fast_forward00:11:31 - um, I mean, empirically that they work well, that there are special cases of
  • fast_forward00:11:36 - the formulation that we have.
  • fast_forward00:11:37 - So this is a purely theoretical statement, right?
  • fast_forward00:11:41 - Algorithms 1 to 10 are special cases of the firmware. Right.
  • fast_forward00:11:44 - And algorithms 1 to 10 had done some effort, fortunately for us,
  • fast_forward00:11:47 - already to show that you're computationally okay.
  • fast_forward00:11:52 - So that's one thing. The other thing is that we could also derive one algorithm,
  • fast_forward00:11:57 - which is the model-free reinforcement learning algorithm, that I find conceptually
  • fast_forward00:12:06 - quite interesting and that we could also directly compare it to other algorithms.
  • fast_forward00:12:10 - If we now first look at the first part of this argument where you say,
  • fast_forward00:12:14 - well, we had the most fundamental formulation of a solution to this set of problems.
  • fast_forward00:12:21 - What are the key ingredients of that solution?
  • fast_forward00:12:25 - I'd say that the key ingredients have already previously been formulated by
  • fast_forward00:12:31 - people like Kappen and Opper and so on.
  • fast_forward00:12:33 - Um the key ingredients is uh to think about optimum control as you know formally
  • fast_forward00:12:40 - it's a minimization problem um but matching two different distributions the
  • fast_forward00:12:47 - one being the control one and the
  • fast_forward00:12:48 - other one being the uncontrolled but conditioned one so these key ideas um,
  • fast_forward00:12:54 - Yeah, have previously actually already been formulated. And I must say it's
  • fast_forward00:12:59 - also not, I mean, you could even go back and talk about Kalman duality in general.
  • fast_forward00:13:05 - Even Kalman already said that the problems of control and the problems of filtering,
  • fast_forward00:13:10 - like state estimation or Bayesian filters, are very similar.
  • fast_forward00:13:16 - And what Bert Kappen did before us, and we now also do, just explicates exactly
  • fast_forward00:13:22 - that relation. and makes it explicit.
  • fast_forward00:13:26 - Now, the difference in the formulation is really that we're talking about processes
  • fast_forward00:13:31 - over state control and not really making any assumptions about the dynamics
  • fast_forward00:13:36 - or control costs or noise in the process,
  • fast_forward00:13:41 - whereas the previous formulations, they discussed processes on the state and
  • fast_forward00:13:46 - for that reason had to do some special assumptions to actually get the theory properly.
  • fast_forward00:13:53 - So in some sense, you had a more reduced formulation of the problem.
  • fast_forward00:13:57 - You just left out a certain number of aspects that you saw as being irrelevant
  • fast_forward00:14:00 - really to the solution. You could put it like this.
  • fast_forward00:14:03 - So how should I imagine the solution that you came up with for these kinds of problems?
  • fast_forward00:14:09 - How does this problem solver really operate?
  • fast_forward00:14:13 - I think most concretely, you could imagine it in terms of that policy that I
  • fast_forward00:14:19 - described, which is actually described by Boltzmann energy.
  • fast_forward00:14:24 - So you actually have a policy which is represented by a Boltzmann energy,
  • fast_forward00:14:30 - which is not really any assumption about the policy, because any conditional
  • fast_forward00:14:34 - distribution can be represented as a Boltzmann distribution.
  • fast_forward00:14:40 - Well, and then the iterative solutions that we proposed translate to iterative
  • fast_forward00:14:45 - updates of that Boltzmann energy.
  • fast_forward00:14:49 - So concretely, I should say that these iterative equations are actually equations
  • fast_forward00:14:55 - which perform that Karpig-Leibner minimization that originally the theory proposes should be done.
  • fast_forward00:15:03 - So these updates of the Boltzmann energy are concretely the argument that has
  • fast_forward00:15:07 - to be implemented eventually.
  • fast_forward00:15:10 - Right, but now basically it means I have a set of policies over which I want to optimize, right?
  • fast_forward00:15:16 - They all have their own level, their Boltzmann energy attached to it,
  • fast_forward00:15:19 - right? And now you're going to update these iteratively.
  • fast_forward00:15:24 - So that means that you are sampling from some set of states that might describe
  • fast_forward00:15:29 - the task that this policy has to be applied to.
  • fast_forward00:15:32 - So to get this iterative function to work reliably, what should be the properties
  • fast_forward00:15:38 - of the states I'm sampling over?
  • fast_forward00:15:40 - Can I just follow any distribution or should it be very regular?
  • fast_forward00:15:45 - So first I want to say that we have a set of policies, which is one,
  • fast_forward00:15:48 - but being described by a distribution.
  • fast_forward00:15:52 - In the exact update where we assume we have a model, we update that Boltzmann
  • fast_forward00:15:59 - energy over the whole space of points.
  • fast_forward00:16:02 - So the whole function we update for all x, for all states.
  • fast_forward00:16:08 - That is the model-based case or the stochastic optimum control case where we assume to have a model.
  • fast_forward00:16:15 - I think what you refer to is when you ask, but for which states do I update
  • fast_forward00:16:20 - it or which states do I have to sample?
  • fast_forward00:16:23 - This is the model-free case, right? Where the system really has to interact with the environment.
  • fast_forward00:16:29 - Well, in that case, the system would have to unroll episodes of experience with the environment.
  • fast_forward00:16:37 - And these episodes of experience give you samples from the process.
  • fast_forward00:16:43 - And you can use these samples of the process as a standard, also with TD and
  • fast_forward00:16:48 - Q-learning and so forth, to do the necessary update of the Boltzmann energy in a stochastic way.
  • fast_forward00:16:54 - So with a learning rate alpha instead of doing the exact update of the energy. Okay.
  • fast_forward00:17:00 - So the point is, what I'm looking for, what are the boundaries on that solution?
  • fast_forward00:17:06 - Not only to formulate the solution, but then to prove that it actually will
  • fast_forward00:17:10 - work and converge, you do have to make assumptions about the space in which
  • fast_forward00:17:14 - this algorithm operates.
  • fast_forward00:17:16 - So what would be these limiting factors?
  • fast_forward00:17:19 - The first one is we only considered observable environments.
  • fast_forward00:17:24 - There is no partial observability in what we discussed.
  • fast_forward00:17:28 - I wouldn't really know yet how it would transfer to partially observable environments.
  • fast_forward00:17:35 - The second thing I think in terms of the,
  • fast_forward00:17:40 - When we can do these updates exactly, I think we prove convergence without any
  • fast_forward00:17:47 - further assumptions, right?
  • fast_forward00:17:49 - So these iterates, they converge.
  • fast_forward00:17:52 - We prove convergence without further
  • fast_forward00:17:54 - assumptions. But this is true if the updates are exact of the energy.
  • fast_forward00:18:01 - So under what circumstances can it be exact? It can be exact if the state and
  • fast_forward00:18:06 - control space is discrete, because then we can represent the energy just as
  • fast_forward00:18:10 - tables over discrete states and actions.
  • fast_forward00:18:14 - In that case, it will converge, no assumptions.
  • fast_forward00:18:19 - If the state space is continuous or otherwise has to be encoded more compactly,
  • fast_forward00:18:27 - the convergence is more difficult.
  • fast_forward00:18:33 - If in the continuous case actually the dynamics is
  • fast_forward00:18:35 - so simple for instance just linear and credit costs um then you can just make
  • fast_forward00:18:41 - a quadratic assumption about that energy and it's almost like the recut equations
  • fast_forward00:18:44 - and so forth you can do the updates exactly and again um it will work um if
  • fast_forward00:18:50 - the dynamics is non-linear well you would have to use a function approximation
  • fast_forward00:18:53 - to represent extended energy.
  • fast_forward00:18:55 - And at that point, exactly, we lose that strict convergence proof and can only
  • fast_forward00:19:02 - hope really, as so often with reinforcement learning and the function approximation.
  • fast_forward00:19:07 - Right. So then the other thing, and sorry if you want, a trick that you applied in this approach,
  • fast_forward00:19:12 - that you in some sense, let's say, completely fragmented your value function,
  • fast_forward00:19:17 - function right because you now added a variable to your system that was locally linked to every state,
  • fast_forward00:19:25 - that was it's not a reward function that relates to that state and if i understood
  • fast_forward00:19:31 - it correctly these reward functions do have to satisfy certain regularities
  • fast_forward00:19:35 - for for this system to operate correctly or or not what what what's the price
  • fast_forward00:19:39 - you pay by distributing your value function in this way hmm.
  • fast_forward00:19:47 - I'm not sure. I mean, these reward functions, we didn't really put assumptions
  • fast_forward00:19:51 - on them, except perhaps for a convergence that they are bounded, as also for Q-learning.
  • fast_forward00:20:00 - So to prove convergence of Q-learning, you would have assumed that they're bounded,
  • fast_forward00:20:04 - but otherwise there's not really constraints.
  • fast_forward00:20:07 - But the point was that in some sense, also contrasting again with this Bellman
  • fast_forward00:20:11 - perspective, where you would have an explicitly defined value function.
  • fast_forward00:20:15 - In this case, you have a more implicitly defined, it seems, value function.
  • fast_forward00:20:21 - Or not. I wouldn't call it value function. We have the rewards and we have the
  • fast_forward00:20:25 - Boltzmann distribution.
  • fast_forward00:20:26 - Nothing more. But you could argue that in iterating these distributed value
  • fast_forward00:20:32 - states, you are approximating something like a value function.
  • fast_forward00:20:37 - Function. In iterating the update of the Bellman, not the Bellman,
  • fast_forward00:20:42 - sorry, the Boltzmann energy.
  • fast_forward00:20:48 - So what I'm after is just to say that you seem to draw a distinction between
  • fast_forward00:20:54 - a value-based approach and your approach where you say it's not value-based in some sense.
  • fast_forward00:21:00 - What I'm saying, well, but maybe implicitly you are value-based,
  • fast_forward00:21:03 - but now you just have sort of hidden it more by putting a distributor in the
  • fast_forward00:21:08 - system linked to every single state. Yeah.
  • fast_forward00:21:12 - So first, yes, the relations in the end become very close. I'll elaborate in a second.
  • fast_forward00:21:19 - But the distinction is not so much between, or I didn't mean to initially distinguish
  • fast_forward00:21:23 - between a value-based approach and our approach, but more like an approach which
  • fast_forward00:21:27 - computes value functions by iterating back the Bellman equation versus by computing other things.
  • fast_forward00:21:34 - In terms of probabilistic inference, right?
  • fast_forward00:21:36 - So the Boltzmann energies are eventually computed minimizing couple of club
  • fast_forward00:21:42 - letter divergences. Fine. Okay.
  • fast_forward00:21:45 - In the Model 3 case, I derived that one equation like that, the Model 3 reinforcement learning equation.
  • fast_forward00:21:51 - And you can analyze the fixed point properties of that equation.
  • fast_forward00:21:55 - I should emphasize it's the fixed point properties, right? It's not the transient
  • fast_forward00:21:59 - of the learning process itself.
  • fast_forward00:22:01 - In that fixed point, this Boltzmann energy has very interesting properties.
  • fast_forward00:22:06 - It's set for non-optimal actions. It goes to minus infinity,
  • fast_forward00:22:09 - making these actions very improbable to be chosen in terms of action selection.
  • fast_forward00:22:15 - For the other ones, it turns out in the fixed point, for optimal actions,
  • fast_forward00:22:19 - it corresponds to the optimal value function, the Boltzmann entity,
  • fast_forward00:22:22 - which is surprising, but it's actually a simple outcome of analyzing the fixed
  • fast_forward00:22:29 - point equation of the update.
  • fast_forward00:22:30 - But now the other thing that was interesting in the approach you described is
  • fast_forward00:22:34 - that you very strongly relied on the inference component also with respect to
  • fast_forward00:22:42 - sort of looking at the consequences of future states that the system might anticipate onto its priors.
  • fast_forward00:22:49 - Right that you sort of as you described yourself as if you just imagine a future
  • fast_forward00:22:55 - goal and then suddenly you just pretend actually you've achieved it you look
  • fast_forward00:22:58 - at the consequence that has on the system um.
  • fast_forward00:23:03 - So what's that dynamic exactly and why did you approach it in these terms?
  • fast_forward00:23:10 - I mean, I don't quite understand what you mean by dynamic.
  • fast_forward00:23:14 - I mean, it's a bit like… Well, it's dynamic because I have to now imagine a future state.
  • fast_forward00:23:19 - I may now include it in my priors and now I can make decisions on that basis
  • fast_forward00:23:23 - so I can propagate myself now
  • fast_forward00:23:25 - to future new events, right? And I can go through that same loop again.
  • fast_forward00:23:30 - Right. Right, so you imagine future events as happening, right?
  • fast_forward00:23:36 - You condition on them and compute a posterior.
  • fast_forward00:23:38 - And in the last framework, yeah, you do iterate by using that posterior again
  • fast_forward00:23:42 - as a prior and condition that again on the same future event and again compute a posterior.
  • fast_forward00:23:51 - Right, I mean... But do you consider that problematic or not? I don't...
  • fast_forward00:23:58 - Algorithmic-wise, I don't see why it's problematic. Do you mean problematic
  • fast_forward00:24:02 - in terms of interpreting how, I don't know, living systems are doing it?
  • fast_forward00:24:06 - Well, to me it seems that this works because you have conditioned it on a very
  • fast_forward00:24:13 - specific definition of the problem domain.
  • fast_forward00:24:17 - So for instance, that I can only have single goals existing at any one point in time, as an example.
  • fast_forward00:24:25 - Oh, no. Why not? No, no. You can arbitrarily condition your future.
  • fast_forward00:24:30 - The future can be represented in a factored way, there can be goals in different
  • fast_forward00:24:36 - representations at different points in time.
  • fast_forward00:24:38 - But wait, if you now propagate that back into your priors, you might have,
  • fast_forward00:24:43 - let's say, unexpected conflicts between these goals that you do have to resolve now in some way.
  • fast_forward00:24:51 - So inferencing graphical models, if there is evidence in a graphical model which
  • fast_forward00:24:56 - are inconsistent, the likelihood is just zero, of that thing,
  • fast_forward00:25:01 - and the messages diverge.
  • fast_forward00:25:02 - And that might, in fact, happen.
  • fast_forward00:25:04 - And that happens in graphical models if there is almost deterministic dependencies,
  • fast_forward00:25:10 - and you get these really inconsistencies.
  • fast_forward00:25:13 - So that's the case when you condition on variables which are observations which
  • fast_forward00:25:17 - are totally inconsistent.
  • fast_forward00:25:19 - If there is a chance, a slight chance of consistency, inference will exactly,
  • fast_forward00:25:28 - generate a compromise. That's the point of it, right? Yes.
  • fast_forward00:25:32 - So, yeah, if you have deterministic models of our future and we condition too
  • fast_forward00:25:37 - many things, it will diverge in a sense.
  • fast_forward00:25:40 - Otherwise, it just behaves as inference in graphical models.
  • fast_forward00:25:43 - Right. But then, so in some sense, that means I just get points in this graph
  • fast_forward00:25:50 - structure that have lost their validity.
  • fast_forward00:25:54 - I do not consider them anymore and I rely on others to make my decisions.
  • fast_forward00:26:03 - Don't know so I mean inferencing graphical models wouldn't really do that right
  • fast_forward00:26:06 - it would find a compromise in a probabilistic sense,
  • fast_forward00:26:10 - another constraint I was worried about is just time like for instance you also
  • fast_forward00:26:15 - mentioned that you have an interest in mapping this to robotics and if you deal
  • fast_forward00:26:21 - with robots the key thing is real world real time so for instance,
  • fast_forward00:26:26 - goal setting might also evolve over different time windows, right?
  • fast_forward00:26:31 - Or I might get, let's say, goal interrupts because the world interferes with
  • fast_forward00:26:35 - my own ideal planning world.
  • fast_forward00:26:37 - So I was wondering how that, how your solution would take these kinds of,
  • fast_forward00:26:43 - let's call them exceptions, into account, or how robust it would be in the face of that.
  • fast_forward00:26:49 - So we don't, I don't know if you have so much experience.
  • fast_forward00:26:53 - Certainly we used that planning as inference method for models that are related
  • fast_forward00:26:58 - to relational reinforcement learning. I didn't talk about this yet.
  • fast_forward00:27:03 - And these are relational.
  • fast_forward00:27:07 - Models on a symbolic level and can describe problems like should I grasp this
  • fast_forward00:27:14 - object or manipulate this object to achieve a goal or to build a tower or things like that.
  • fast_forward00:27:18 - Do I first have to open a door before I get an object and all these kind of things.
  • fast_forward00:27:23 - So we use these models and planning as inference in these kind of models which is really fast.
  • fast_forward00:27:29 - Really, because it's on a symbolic level. Very fast compared to all the low-level motion stuff.
  • fast_forward00:27:35 - So in those cases, what you call interrupts or unexpected events or things like
  • fast_forward00:27:40 - that, would just require to sort of redo the whole inference,
  • fast_forward00:27:46 - which, because it's on a symbolic level, is really very quick.
  • fast_forward00:27:50 - So, I don't know, currently in practice in robotics, I'd say the inference is
  • fast_forward00:27:56 - so fast that you just update online all the time.
  • fast_forward00:28:01 - Okay, so you're saying this will just be equalized out by compute speed.
  • fast_forward00:28:06 - You could put it like that, yeah. By just having the system to just update itself
  • fast_forward00:28:10 - and recompute the posterior whenever some new evidence comes in.
  • fast_forward00:28:16 - Right. So then also at some point in your presentation, you gave a little scheme
  • fast_forward00:28:20 - where you distinguished different variations of problem-solving approaches,
  • fast_forward00:28:26 - like model-based versus model-free.
  • fast_forward00:28:31 - Oh, that one. So what's that structure that you had exactly in mind there?
  • fast_forward00:28:35 - How does it organize the different approaches that are around?
  • fast_forward00:28:40 - Um i didn't really i didn't really think that there is like one structure which
  • fast_forward00:28:46 - combines all these approaches uh i mean that that diagram was just to to sort
  • fast_forward00:28:51 - of give an overview of what,
  • fast_forward00:28:53 - what kind of approaches people follow in reinforcement learning, right?
  • fast_forward00:28:56 - So model-based, the path going over first learning transition models and model-free,
  • fast_forward00:29:00 - just learning value functions.
  • fast_forward00:29:03 - I mean, you know Dyna, which combines the two in one framework.
  • fast_forward00:29:08 - I don't know. So our own work is mostly, actually, the work that we use in robotics,
  • fast_forward00:29:14 - which I didn't talk much about today, is really fully model-based.
  • fast_forward00:29:18 - So we always actually follow model-based approaches.
  • fast_forward00:29:21 - And it's only in what I talked about today, in that more recent work,
  • fast_forward00:29:26 - that we came up with a model-free reinforcement learning algorithm,
  • fast_forward00:29:28 - which I find interesting for other reasons.
  • fast_forward00:29:32 - But in robotics, I must say, I'm more a model-based proponent.
  • fast_forward00:29:36 - Why is that? Why do you find that more interesting or relevant?
  • fast_forward00:29:40 - In particular, for the types of problems that we did in relational reinforcement
  • fast_forward00:29:45 - learning, where the state space is inherently exponential in the number of objects.
  • fast_forward00:29:50 - So the state is described by all the relations between objects.
  • fast_forward00:29:54 - So if you have 10 objects, there is on the order of 10 squared binary relations between them.
  • fast_forward00:30:01 - All of them can have a value, so your state space is, say, 2, 2, 2, you know.
  • fast_forward00:30:09 - So the state space is exponential in those cases and how do you do learning
  • fast_forward00:30:13 - in these kinds of spaces? Um, and, and.
  • fast_forward00:30:17 - There is model-free approaches as well, who sort of represent also the policy
  • fast_forward00:30:24 - directly on some relational features, so first-order logic features,
  • fast_forward00:30:28 - and then use policy gradients to actually optimize them.
  • fast_forward00:30:31 - But we think that the generalization is actually much stronger if you try to
  • fast_forward00:30:35 - learn models from the experiences.
  • fast_forward00:30:37 - And the models in those cases are represented by stochastic relational rules.
  • fast_forward00:30:42 - Almost a bit like STRIP, but there is stochastic in their first order.
  • fast_forward00:30:46 - And it's not our own work. We use that work to learn these rules.
  • fast_forward00:30:51 - And we find it fascinating how strongly they generalize. So from only a couple
  • fast_forward00:30:55 - of experiences that something rolls when you push it, the robot generalizes
  • fast_forward00:30:59 - quite, I don't know, rationally one could say, I don't know,
  • fast_forward00:31:03 - quite interestingly to other objects and so forth.
  • fast_forward00:31:05 - So it is what I find interesting about model-based approaches
  • fast_forward00:31:08 - is the ability to generalize experience and but
  • fast_forward00:31:12 - the problem of model-based of course is well now you have that
  • fast_forward00:31:14 - nice model but how do you translate it to actions and this
  • fast_forward00:31:17 - is where then the planning is inference right exactly but to
  • fast_forward00:31:20 - what extent are these models uh actively acquired uh nice work so the the last
  • fast_forward00:31:28 - two years or so we've been um actually investigating a lot in in active exploration
  • fast_forward00:31:34 - or active learning so where a system should you know decide on actions that
  • fast_forward00:31:38 - maximize information gain.
  • fast_forward00:31:41 - And there's the fear of active learning in machine learning, right?
  • fast_forward00:31:45 - So one of the questions we had is how can we sort of transfer these methods,
  • fast_forward00:31:50 - the existing methods, on the relational for relational reinforcement learning.
  • fast_forward00:31:54 - And that requires that machine.
  • fast_forward00:31:58 - Estimating information gain or estimating your uncertainty of predictions,
  • fast_forward00:32:02 - as it is done also implicitly in R-Max or the Bayesian Exploration Bonus,
  • fast_forward00:32:06 - how to do the same thing with stochastic rules.
  • fast_forward00:32:10 - We did that. We call this relational exploration. That means that if the robot
  • fast_forward00:32:16 - has observed a ball rolling,
  • fast_forward00:32:18 - a green ball rolling and then a blue ball rolling, then it It maybe would not
  • fast_forward00:32:22 - be interested anymore in another yellow ball, but instead try to roll something
  • fast_forward00:32:28 - which looks different to a ball.
  • fast_forward00:32:29 - So absolutely. So these are exploration strategies that we investigated and
  • fast_forward00:32:35 - which are really important in these exponential state spaces to lead to efficient learning.
  • fast_forward00:32:40 - But for you, the main difference is that you say, look, if this state space
  • fast_forward00:32:44 - gets too large, you just need a better strategy to map it.
  • fast_forward00:32:49 - And that's an active learning. learning but or an active exploration component
  • fast_forward00:32:54 - while um intrinsically intrinsically you don't need to change your whole approach
  • fast_forward00:33:01 - it's just the way how you explore that state space i i agree so,
  • fast_forward00:33:05 - intrinsically uh it's a funny choice of word because of intrinsic motivation
  • fast_forward00:33:09 - but um it's the same principle so it's still the same principle of maximizing
  • fast_forward00:33:14 - information gain uh which is approximated as it is in Rmax and our vision exploration bonus.
  • fast_forward00:33:20 - What is necessary is to transfer that same intrinsic principle to other types
  • fast_forward00:33:26 - of representations, namely those relational representations or these stochastic rules we had.
  • fast_forward00:33:31 - Would you think this game would change a lot when I would impose a memory capacity constraint?
  • fast_forward00:33:40 - Um don't know um you mean
  • fast_forward00:33:44 - the exploration strategy would change
  • fast_forward00:33:47 - yeah or or the model the whole model construction
  • fast_forward00:33:50 - phase well in the model construction there is a regularization penalizing size
  • fast_forward00:33:54 - of the model so putting a hard limit on the model sometimes can be even shown
  • fast_forward00:34:01 - to be sort of dual to actually regularization but um yeah i'm not sure if it
  • fast_forward00:34:06 - would change So why this is interesting to me is, look, I try to understand how the brain works.
  • fast_forward00:34:09 - And the brain just doesn't have these luxuries of, you know,
  • fast_forward00:34:14 - infinite memory or infinitely fast processing speeds and so on.
  • fast_forward00:34:20 - And so this, of course, raises this issue that maybe the brain is operating
  • fast_forward00:34:25 - in a part of problem-solving state space where we just are not really looking
  • fast_forward00:34:30 - with our algorithms because we're not looking at the same constraints.
  • fast_forward00:34:34 - So it's for that I'm asking if I take a constraint not as a memory capacity,
  • fast_forward00:34:38 - would that change the game from your perspective.
  • fast_forward00:34:43 - I'd say I don't I don't know if I can,
  • fast_forward00:34:48 - say anything to that. I have the impression that the kind of models that we
  • fast_forward00:34:52 - learn are so small compared to the things that humans really learn.
  • fast_forward00:34:57 - Humans would have much more capacity than what our models that use these stochastic
  • fast_forward00:35:02 - relational rules, for instance.
  • fast_forward00:35:05 - Actually, I find it sometimes quite surprising how much humans can memorize,
  • fast_forward00:35:09 - children especially, without actually abstracting just as if they would just
  • fast_forward00:35:14 - throw it away before they actually later on. Without even being programmed by you. It's amazing.
  • fast_forward00:35:19 - Yes. So the other thing here is also, for instance, a lot of these methods that
  • fast_forward00:35:27 - are very advanced and you guys have a deep understanding of them,
  • fast_forward00:35:30 - but they are very often predicated on fairly strong assumptions.
  • fast_forward00:35:34 - For instance, one thing that I find often worrisome is that,
  • fast_forward00:35:39 - for instance, people sort of
  • fast_forward00:35:41 - very loosely say, well, okay, let's assume I have defined my state space.
  • fast_forward00:35:44 - And now on the basis of that I'm going to show you all sorts of properties of
  • fast_forward00:35:50 - my algorithm or I'm going to learn a model I'm going to do problems with and so on.
  • fast_forward00:35:54 - Isn't that assumption actually too strong? Isn't it?
  • fast_forward00:35:57 - Shouldn't we also think more about policy learning together with really the
  • fast_forward00:36:01 - state space acquisition from experience?
  • fast_forward00:36:05 - Yeah. Totally agree.
  • fast_forward00:36:08 - Yes, I totally agree. I mean, the question of where the representations comes
  • fast_forward00:36:12 - from is absolutely fundamental.
  • fast_forward00:36:15 - Maybe there are two ways to go about this.
  • fast_forward00:36:19 - The one is really having the ambition that the system should really invent its
  • fast_forward00:36:22 - own notions of state, its own representations and everything from scratch,
  • fast_forward00:36:27 - which has been the ambition,
  • fast_forward00:36:30 - for some while by researchers, and ideally even under partial observability and so forth.
  • fast_forward00:36:37 - Which, even myself, I was thinking about that. But more and more,
  • fast_forward00:36:41 - I think this is very tough.
  • fast_forward00:36:43 - Still, we should still continue trying, right? The more I actually work with robotics,
  • fast_forward00:36:48 - the more I have the impression that it might be worth before trying to invent
  • fast_forward00:36:54 - algorithms that can invent representations for all possible worlds,
  • fast_forward00:36:59 - to actually understand that actually our world, our 3D world, is quite special.
  • fast_forward00:37:07 - And it's special in the sense that it's 3d there is
  • fast_forward00:37:10 - physics a lot of things in our world are actually about rigid bodies and it's
  • fast_forward00:37:16 - becoming now very robotics talk right there's a lot of things about really kinematics
  • fast_forward00:37:22 - so there is degrees of freedom in our environment that you can manipulate and
  • fast_forward00:37:25 - so forth so our actually true world is very very structured,
  • fast_forward00:37:30 - So in that sense, why not as a simple step, because we're limited as researchers,
  • fast_forward00:37:36 - as humans, first try to understand those particular structures that are inherent
  • fast_forward00:37:40 - in our world and think about what would be representations to actually deal with those.
  • fast_forward00:37:45 - So that's also a more robotics-like, pragmatic approach.
  • fast_forward00:37:49 - But it's also funny that you're sort of following a little bit your physics
  • fast_forward00:37:52 - training or intuitions again by saying, oh no, let's make the spherical cow
  • fast_forward00:37:57 - and then we start from there. Frankly, I'm not sure.
  • fast_forward00:38:01 - I think it's the opposite in physics and training. I think the typical physicist
  • fast_forward00:38:05 - would go the first approach because he wants to be always general and work in
  • fast_forward00:38:08 - every possible world and derive optimal solutions in every possible world.
  • fast_forward00:38:11 - And in a sense, also machine learning and optimal control and MDPs, they do that, right?
  • fast_forward00:38:18 - They define problems which are seemingly total generally because you can describe
  • fast_forward00:38:24 - everything as an MDP or a control problem, right?
  • fast_forward00:38:26 - And you can embed everything in a vector space, which is sort of true.
  • fast_forward00:38:30 - But it might neglect the fact.
  • fast_forward00:38:33 - The actual problems we are faced with are very structured.
  • fast_forward00:38:38 - And acknowledging that specific structure is maybe less a theoretical physicist
  • fast_forward00:38:43 - thing, but more really in engineering and robotics.
  • fast_forward00:38:46 - But do you feel that there are already solutions of that kind on the table? No.
  • fast_forward00:38:53 - Looking more and more into robotics, I think that there is, of course, there is engineering,
  • fast_forward00:39:02 - which means that the people who program the roboticists, they have these concepts
  • fast_forward00:39:07 - and these specific representations and understanding of kinematics of worlds and so forth.
  • fast_forward00:39:12 - And from their understanding, then program the robot to deal with that.
  • fast_forward00:39:15 - But I do not have the impression that the machines themselves,
  • fast_forward00:39:20 - so let's say the inference mechanism in the models that I talked about today,
  • fast_forward00:39:25 - that they would actually do inferences about true physical situations.
  • fast_forward00:39:30 - Right. Or that we would have inference machines which could do inferences about
  • fast_forward00:39:34 - what is a stable physical situation, like probabilistic inference over that,
  • fast_forward00:39:38 - or where might be a degree of freedom in the world.
  • fast_forward00:39:43 - So doing inferences in these spaces, they're so structured that I think we do
  • fast_forward00:39:47 - not yet know how we could do probabilistic inference in those spaces.
  • fast_forward00:39:50 - Right. So what would be a good benchmark for these kinds of models?
  • fast_forward00:39:56 - Hmm...
  • fast_forward00:39:59 - Okay, so there is the old Köhler, Wolfgang Köhler, right?
  • fast_forward00:40:04 - He was a psychologist in the beginning of the 20th century. He did these experiments with monkeys.
  • fast_forward00:40:12 - Japanese, yeah. Exactly. I loved them on which island? I forgot.
  • fast_forward00:40:19 - So he has this book, which is called Intelligence Tests of Apes, something like that.
  • fast_forward00:40:23 - In these books he describes actually very nice behaviors which I would call
  • fast_forward00:40:29 - really goal directed behaviors and one of these behaviors is where there is
  • fast_forward00:40:34 - a banana at the ceiling and the robot is trying to get it but it's too high
  • fast_forward00:40:37 - it's a chimpanzee who tries to get it not a robot oh sorry,
  • fast_forward00:40:41 - you see where I'm getting at right so yeah the ape wants to get it it's too
  • fast_forward00:40:48 - high he jumps and for five minutes doesn't achieve it,
  • fast_forward00:40:51 - gets sort of depressed at least that's the storytelling that Köhler actually
  • fast_forward00:40:56 - does in the book sits in the corner.
  • fast_forward00:40:59 - Uh sort of depressed for a while and then suddenly uh
  • fast_forward00:41:02 - points his eyes towards the banana and then points his
  • fast_forward00:41:05 - eyes at the box in a corner and then again at
  • fast_forward00:41:08 - the banana in the box and then jumps up grabs the box goes on top of the box
  • fast_forward00:41:12 - and gets it um and i i wish you know robots could do that and i i think if i
  • fast_forward00:41:19 - would be psychologist so i would be very interested to to actually model what's
  • fast_forward00:41:24 - going on in that head of the monkey.
  • fast_forward00:41:29 - Because I think he did inference in a very structured way, inference about our physical world.
  • fast_forward00:41:35 - He did inference about, let me pull that box here and get on top of that.
  • fast_forward00:41:39 - And it's exploiting physics very, very much.
  • fast_forward00:41:43 - Yeah. We had a beautiful lecture on this last week by Alex Kacelnik,
  • fast_forward00:41:46 - exactly on this famous experiment by Koehler with also many other experiments
  • fast_forward00:41:52 - in that domain. So you would follow, you would stick to such a benchmark,
  • fast_forward00:41:56 - you'd be happy with that.
  • fast_forward00:41:57 - Absolutely. If the robots would do that, I'd be… Very good.
  • fast_forward00:42:00 - But now, I don't know whether you guys ever apply your own methods in an autobiographical way.
  • fast_forward00:42:05 - Because you could also say, look, every machine learning person is trying to
  • fast_forward00:42:09 - optimize their own policy in this complex world.
  • fast_forward00:42:13 - And in some sense, we also would like to make statements about biological systems,
  • fast_forward00:42:17 - about physical, psychological systems, and so on using these methods.
  • fast_forward00:42:20 - But there might be a possibility that, of course, you guys have optimized yourself
  • fast_forward00:42:25 - in some sort of local minima in this larger space of all possible policies and
  • fast_forward00:42:29 - problem-solving solutions,
  • fast_forward00:42:30 - that actually the links with natural systems has gotten lost.
  • fast_forward00:42:36 - Okay. So where do you stand on that?
  • fast_forward00:42:40 - To what extent also the models we now very superficially discussed now in this
  • fast_forward00:42:44 - interview, I mean, where is really the leverage, right?
  • fast_forward00:42:48 - Because, for instance, we all know the big hype that has been going around and
  • fast_forward00:42:53 - the big enthusiasm and hope for the last 30 years on these kinds of methods.
  • fast_forward00:42:58 - And sometimes when we look back, we can also say, well, yeah,
  • fast_forward00:43:00 - okay, it kept a lot of people busy. That's all very positive.
  • fast_forward00:43:03 - But did we really sort of make a huge step forward in understanding biological
  • fast_forward00:43:07 - systems or psychological systems?
  • fast_forward00:43:09 - And there, it's not so clear. Yeah, I agree that there is not necessarily always
  • fast_forward00:43:15 - links to the biological system.
  • fast_forward00:43:18 - Of the methods. And I would even claim that's very often not the goal to keep
  • fast_forward00:43:25 - these links to the biological paradigm.
  • fast_forward00:43:29 - Why not? It can be one goal for some people, but it's also an engineering goal
  • fast_forward00:43:37 - really to just design systems which do things good in a good way and optimize something or whatever.
  • fast_forward00:43:44 - But let me put it the other way.
  • fast_forward00:43:45 - I think in order to understand living systems, it is also a good idea to,
  • fast_forward00:43:52 - to understand our world, our environment, and to understand what it means to
  • fast_forward00:43:57 - behave, to organize behavior in that environment.
  • fast_forward00:44:01 - Totally, just on that level, I tend to say functional also, without caring for the substrate,
  • fast_forward00:44:08 - without caring for, I don't know, the constraints, the concrete biological constraints
  • fast_forward00:44:13 - that there are, but just to understand what is actually the structure of problems
  • fast_forward00:44:16 - that things in the world are being faced with, no matter if these are robots or not or other systems.
  • fast_forward00:44:23 - And I think that the engineering methods or robotics can contribute on that side.
  • fast_forward00:44:28 - And I also think that, well, it's important to understand actually the problems
  • fast_forward00:44:35 - when you then would go back and ask how could biological systems actually face these problems.
  • fast_forward00:44:39 - So it's a bit more a normative perspective then.
  • fast_forward00:44:43 - Like this is what the system should do.
  • fast_forward00:44:47 - Should do is even too much. what the system is confronted with,
  • fast_forward00:44:51 - understand the structure of what it is confronted with.
  • fast_forward00:44:53 - But now, if we look at nature or psychology, we might see that actually notions
  • fast_forward00:44:59 - of optimality might not hold so well, right?
  • fast_forward00:45:02 - Like human decision making is in many occasions suboptimal.
  • fast_forward00:45:06 - I'm here wasting your time and you're still being polite so that you could also
  • fast_forward00:45:10 - consider that suboptimal, right?
  • fast_forward00:45:12 - So in the face of let's say psychological reality and behavioral reality,
  • fast_forward00:45:18 - how do these notions of optimality actually hold up to that?
  • fast_forward00:45:22 - Optimality, so I have a very pragmatic actually relation to optimality I think,
  • fast_forward00:45:30 - many empirical things can be described as if they would adhere optimality principles,
  • fast_forward00:45:36 - it's a misunderstanding to actually say that this would have any meaning like
  • fast_forward00:45:41 - if it wants to be optimal or whatever let me take an example,
  • fast_forward00:45:45 - I mean the trajectory of a particle in physics, right? You can say it minimizes action.
  • fast_forward00:45:52 - Or some kind of field in physics behaves as if it would minimize an energy.
  • fast_forward00:45:58 - That energy or that action is only a scientific concept that we scientists have
  • fast_forward00:46:03 - devised to describe the actual system there.
  • fast_forward00:46:10 - Cost functions or optimality objectives like these things.
  • fast_forward00:46:14 - These are just designed by us scientists to describe systems in the world.
  • fast_forward00:46:20 - It's just a means to describe systems in the world, to say the system behaves as if it is optimal.
  • fast_forward00:46:24 - So I think it's a misunderstanding to really think that we believe machines must be optimal.
  • fast_forward00:46:31 - It's just a scientific tool to describe them as being optimal with respect to
  • fast_forward00:46:35 - some arbitrary cost function.
  • fast_forward00:46:36 - But for instance, humans have many biases in their decision-making,
  • fast_forward00:46:40 - right? And there's some beautiful books written about that.
  • fast_forward00:46:43 - One might be, for instance, many people in gambling believe they have much more
  • fast_forward00:46:46 - control over the outcome of a bet than they really have.
  • fast_forward00:46:50 - So, you see, in psychological reality, you might not be able to really optimize
  • fast_forward00:46:55 - so cleanly on some goal function, even though that is really at the core of
  • fast_forward00:46:59 - the methods that you're developing.
  • fast_forward00:47:01 - No, I think, and this is almost a trivial statement, that any behavior can be
  • fast_forward00:47:06 - described as being optimal. It's just a question of optimal with respect to what.
  • fast_forward00:47:09 - And I mean this, and it's no more statement when I'm saying I'm interested in
  • fast_forward00:47:16 - describing algorithms that can do optimal things.
  • fast_forward00:47:19 - It's only that I think that this is an abstraction of how to describe behavior.
  • fast_forward00:47:24 - I think even what people call sub-rational behavior or limited behavior,
  • fast_forward00:47:28 - of course you can describe it as being optimal with respect to something, some other objective.
  • fast_forward00:47:33 - Okay, but then, of course, there's another risk that's looming,
  • fast_forward00:47:37 - another challenge that I could say, yeah, but then, in that case,
  • fast_forward00:47:40 - your method is super powerful because it can never fail.
  • fast_forward00:47:43 - So then what do we learn? Well, the thing is that you now sort of shifted your
  • fast_forward00:47:48 - level of scientific description.
  • fast_forward00:47:50 - You know, you do not describe….
  • fast_forward00:47:53 - Behavior phenomena phenomena directly anymore in
  • fast_forward00:47:56 - a direct representation in terms of describing the policy but
  • fast_forward00:47:59 - you described you shifted your level of scientific description to the
  • fast_forward00:48:02 - level of describing what they optimize and at
  • fast_forward00:48:05 - first sight it's nothing more than just a shift you didn't gain anything by
  • fast_forward00:48:09 - that but potentially what you could gain is that this other description is more
  • fast_forward00:48:14 - compact and that's really the only scientific gain of that that you can describe
  • fast_forward00:48:19 - things more compactly. And this is how it was in physics always.
  • fast_forward00:48:22 - The reason why you would want to describe particles as minimizing action,
  • fast_forward00:48:26 - so the trajectory of particles, is because it's a very concise description of what's happening.
  • fast_forward00:48:31 - I mean, alternatively, you could just write down for every possible thing what
  • fast_forward00:48:35 - happened, but it's more concise to describe it like that.
  • fast_forward00:48:37 - Same with behavior like human motion.
  • fast_forward00:48:41 - Saying that human motion, as in
  • fast_forward00:48:44 - Walpert's work, behaves as if it would be stochastically optimal control.
  • fast_forward00:48:49 - Don't overstate this. It's not meaning that they want to optimize it or so.
  • fast_forward00:48:54 - It's just a statement that you can shift your scientific description of human
  • fast_forward00:48:58 - behavior onto an abstraction where you say the behavior, the motion is as if
  • fast_forward00:49:05 - it would minimize that objective function.
  • fast_forward00:49:07 - Right. Just a shift of description, nothing more. And a description of it, right?
  • fast_forward00:49:10 - Right. The hope would be that it's so concise that, for instance,
  • fast_forward00:49:13 - with robots, it's easier to specify or be creative in saying what they might
  • fast_forward00:49:20 - want to optimize and then derive behavior from that rather than programming behavior directly.
  • fast_forward00:49:25 - It's a shift of programming languages almost, if you want to say.
  • fast_forward00:49:28 - Now we're programming objective functions instead of policies directly. Exactly.
  • fast_forward00:49:32 - So Mark, so you're deeply involved in developing these sort of also new perspectives
  • fast_forward00:49:39 - on problem solving and machine learning, which also have quite some impact.
  • fast_forward00:49:45 - But so given this experience, what would you see as Mark's law for us to understand reality? Yeah.
  • fast_forward00:49:52 - Ah, I have no idea.
  • fast_forward00:49:55 - Mark's law. Acknowledge the structure. I don't know.
  • fast_forward00:49:59 - That's an important thing, which sometimes I don't feel myself that my own research
  • fast_forward00:50:04 - makes so much progress with that. But actually, I think this would be the most important thing.
  • fast_forward00:50:09 - Acknowledge really the structure of the problems.
  • fast_forward00:50:14 - Because, yeah, that's important. Right. And last point.
  • fast_forward00:50:19 - So five years from now, I'm going to come visit you wherever you are.
  • fast_forward00:50:21 - Maybe Stuttgart and I'm going to confront you with a prediction you're going
  • fast_forward00:50:25 - to make today but I'm going to say look Mark you predicted to me today or five
  • fast_forward00:50:29 - years back that the following would happen now did it really come out like that
  • fast_forward00:50:33 - was your prediction really supported,
  • fast_forward00:50:37 - what's this prediction you would like to make today um that you can test me in five years right,
  • fast_forward00:50:43 - um um well so our goal actually is and I usually say to my students in five
  • fast_forward00:50:50 - years as a point of motivation, that we would have that robot that autonomously
  • fast_forward00:50:55 - explores its environment,
  • fast_forward00:50:57 - and discovers all degrees of freedom.
  • fast_forward00:51:01 - That sounds like a simple statement initially, but if you look around the room
  • fast_forward00:51:05 - right now and think about all degrees of freedom in that room,
  • fast_forward00:51:08 - kinematic ones, there's a lot.
  • fast_forward00:51:10 - And in order to discover them, the robot would have to start pushing,
  • fast_forward00:51:13 - kicking, letting fall, letting drop, I don't know, everything grasping all of
  • fast_forward00:51:19 - these things and shake it and see if something moves. And that's something I would want to...
  • fast_forward00:51:24 - Okay, I'll come and see if the robot tore down your lab five years from now.
  • fast_forward00:51:27 - Yeah. So, Marc Doussaint, thank you very much for this conversation.
  • fast_forward00:51:30 - You're welcome. Thank you.
  • fast_forward00:51:31 - Music.
  • fast_forward00:51:37 - The CSN Podcast was produced by the Convergent Science Network of Biometrics
  • fast_forward00:51:43 - and Biohybrid Systems, a project funded by the European 7th Research Framework Programme.
  • fast_forward00:51:51 - For more interviews, recorded lectures or upcoming conferences in the field
  • fast_forward00:51:56 - of biometrics and biohybrid systems, go to csnnetwork.com.
  • fast_forward00:52:02 - Music.

Be the first to leave a comment

Leave a comment

Your email address will not be published. Required fields are marked *

Convergent PozitivLinia Logo

Exploring the convergence of neuroscience, robotics, and AI through conversations with leading researchers since 2010.

A project of the Convergent Science Network Foundation.

© CSN Podcasts. Developed by IMCreativeWEBC

0%

Login to enjoy full advantages

Please login or subscribe to continue.

Go Premium!

Enjoy the full advantage of the premium access.

Stop following

Unfollow Cancel

Cancel subscription

Are you sure you want to cancel your subscription? You will lose your Premium access and stored playlists.

Go back Confirm cancellation