Thursday, November 16, 2006

Videos - Reinforcement Learning Control Tasks

These videos are from two reinforcement learning control problems I setup for my master's thesis work last year.


Pendulum Swing-Up
A physically-simulated pendulum is controlled by a Verve agent (http://verve-agents.sourceforge.net). Based on simple reinforcement signals (+1 when the pendulum is close to vertical, -1 otherwise), the agent learns to swing the pendulum upright and balance it after about 60 trials.






Cart-Pole/Inverted Pendulum
A physically-simulated cart is controlled by a Verve agent (http://verve-agents.sourceforge.net) in order to balance an attached pole. Based on simple reinforcement signals (-1 when the pole falls over or the cart goes off the edge of the platform, +1 otherwise), the agent learns to balance the pole for 30 minutes after about 600 trials.


Videos - Magnet Toy Simulation

These are videos of a simulation I made earlier this year. I was trying to simulate a magnet toy using OPAL/ODE. The toy has two sets of magnets, each attached to a rotating axle. Here's a picture (thanks, JO):












Videos - Artificial Evolution of Humanoid Behaviors

These videos are from a project I did in 2003. A physically simulated humanoid is controlled by an artificial neural network which senses joint angles and controls muscle forces. A genetic algorithm optimizes the neural network weights to improve performance on some motor control task.


Jumping Vertically





Standing Upright





Walking Forward


Tuesday, November 07, 2006

We are All Allotted a Constant Amount of Mental Energy

Given that we humans all have similar brain structures of comparable sizes, bandwidths, etc., we are all allotted a constant amount of mental energy which enables us to perform a certain amount of mental work per unit time. You could think of "mental work" as "planning" or "thinking." To be more precise, let's define it as reducing the uncertainty (or increasing the predictability) of a situation.

Let's look at some examples:
  1. Visual scenes: The more time you spend looking at a landscape, an interior space, or a painting, the better you are able to predict it. At first you aren't able to predict the spatial relationships between elements, but after some time it becomes easy.
  2. Logic and word puzzles: After expending a certain amount of mental energy, uncertainty is reduced, and the puzzles disappear.
  3. People: At first, strangers can be unpredictable. The more you interact with a person, the better you understand him or her. It becomes easier it is to predict his or her behavior. (Most people's personalities are moving targets, though, so it isn't possible to reach a state of completely accurate prediction.)
  4. TV shows: The more you watch a certain show, the better you are able to predict subtle interactions between characters and high-level plot elements.
  5. Sports and music: Learning new motor behaviors (like new sports, new pieces of music, and new musical instruments) requires a lot of attention. It takes effort to coordinate muscle movements into desired patterns. Over time, though, repetitive training produces smooth, synchronized motions with little effort.
  6. Cooking: The more brain-time you invest, the better you understand how ingredients interact, and the better you can predict what things will taste like.
  7. Investments: At first you have no idea how stable various markets are. Over time, your uncertainty is reduced, and you can make well-informed decisions based on past experience.
In each situation it takes a certain amount of mental work to reduce uncertainty, just like it takes a certain amount of physical work to move heavy objects. You could say that complex situations are "heavier" than others.

Even though most of our brain structures function in parallel, the whole attention system is a serial mechanism. Think about it. You can only focus on one discrete thought at a time. Try looking at a complex scene or object. You are able to look at the whole thing at once, but you can't attend to more than one component at a time. Right now I'm looking at a house plant in my living room, and I can't simultaneously think "plant" and "leaf." I have to focus on either a small part or the whole thing. It's complicated because you can think of similar sets of objects at a time ("all the books on my bookshelf"), but you're still limited to a single discrete thought.

Imagine your attention system constantly switching among various thoughts. It spends some time working on a crossword puzzle, switches for a half second to think about what to have for lunch, switches back to the puzzle, switches to the sound coming from the radio for a bit, switches back to the puzzle... Every time attention shifts, it starts applying mental force to a new mental object, moving it a little bit across the spectrum from unpredictable to predictable.

The point here is that we all have a similar (within an order of magnitude) amount of mental energy available to us per unit time. Since birth we have all expended a similar amount of this energy on something. So everyone must have some kind of hidden talent related to those things that occupy most of his or her thoughts.

(The whole idea of "useful" mental work is another story. Utility can be defined in a variety of ways. A common one might be something along the lines of "the good of the many." If a person's goals are aligned with the utilitarian viewpoint, he or she would spend his or her mental energy on problems that benefit the most people. If the reward hypothesis is correct, we spend our mental energy on those things that we expect [based on previous experience] will bring us the most rewards. This gets into the whole area of motivation, which is beyond the scope of this article. To attain the highest level of utility, however it is defined, it is probably necessary to spend time thinking of ways to improve one's own thinking abilities, or metalearning.)

Now Listening to...

...audio lectures from Jeremy Wolfe's course, Intro to Psychology, Fall 2004, on MIT OpenCourseWare. Even though it's an intro course, I think it'll be good. Jeremy sounds like a great teacher (you can just tell after listening to him for 30 seconds). And I think it's good to hear the fundamentals lots of times from a variety of teachers.

I didn't quite finish to Gerald Schneider's Animal Behavior course (I finished 22 out of 37 lectures). I heard all I wanted to hear and decided to move on.

Update on Verve Development

For a while my posts have been focused on things other than the Verve project itself. One reason is that I enjoy posting about random interesting ideas that pop into my head. The other reason is that I've been thinking of starting a new software development effort. It would have the same general goals as Verve, but it would use much more advanced methods. For instance, I have been doing a lot of research into the general topic of context representation. I have prototyped a few subcomponents and have drawn lots of plans. Rather than ripping out the existing context representation in Verve (i.e. its dynamically-growing radial basis function system) and reworking all the components that connect to it, I'll probably just leave it how it is and continue building this new system.

Things are going very well overall. I'm having a blast. I think it's important to have a blast when you're doing research. It's good for morale. And if your morale is suffering, you're not going to do good research.

I'll post more details as things progress.

Thursday, September 21, 2006

Intuitiveness

What does it mean for something to be "intuitive?"

It must have something to do with our intuition. Ok... so what's "intuition?" I like to define it as previous knowledge (loosely defined), either instinctive or gained through learning. Thus, something that is intuitive is something that takes advantage of previous knowledge. A key point here is that intuitiveness is subjective.

If you have played poker games in the past, the rules of a new poker game will be intuitive if they rely on knowledge gained from other poker games. If you have used Microsoft products in the past, new Microsoft products will be intuitive as long as they are designed like the old ones. If you have driven a car in the past, driving a new car should be easy. (Driving may not be very intuitive at first since we don't otherwise press levers to change velocity and turn a wheel to change directions... except in video games.) In all of these cases standards are important since they ensure that previous knowledge is exploited.

This thought process started a year and a half ago in a class assignment for "Interaction Methods for Emerging Technologies." The assignment was to explain why direct manipulation devices are usually preferred by users. (Direct manipulation devices are those in which the user's actions directly affect the end object, as opposed to devices that add one or more levels of indirect manipulation. For example, a computer mouse has one level of indirect manipulation since it indirectly controls the pointer on the screen.) This was my answer:
Why are direct-manipulation interfaces preferred?

I will use the phrase "knowledge transfer" to refer to the amount of previous knowledge that can be applied to a new domain. Direct-manipulation exploits a lot of knowledge transfer because user's manipulate the device in a similar way to how they manipulate everyday objects. Direct-manipulation usually requires little learning, thus less effort and/or frustration when using a new device.

Additionally, I would say that intuitiveness in any domain is directly proportional to the amount of knowledge transfer being used, maybe going so far as to say intuitiveness is equivalent to the utilization of knowledge transfer. So things can be intuitive to some people and not others depending on their experience. Direct-manipulation is more intuitive to almost everyone because almost everyone has had a lot of experience manipulating everyday objects. A new Microsoft product, on the other hand, would be intuitive for people experienced with Microsoft products because of the knowledge transfer involved, but not for others who aren't used to them (hence the effectiveness of interface standardization).

Monday, February 28, 2005

Monday, September 18, 2006

ASME IDETC & CIE Conference 2006


I presented a paper on Verve at the 2006 ASME (American Society of Mechanical Engineering) IDETC & CIE conference (International Design Engineering Technical Conferences and Computers and Information in Engineering... whew). The official reference is the following:

Streeter, T., Oliver, J., & Sannier, A. 2006. A General Purpose Open Source Reinforcement Learning Toolkit. In Proceedings of the ASME International Design Engineering Technical Conferences and Computers and Information in Engineering Conference.
The presentation went pretty well. I was put in a strange topic area, though (Knowledge Management in Design Automation). I think the audience liked the live demos, and I had some good questions from them afterwards. The paper and presentation slides are available in the publications section of my website.

I attended a talk called "The Spirituality of Engineering" by two professors from Delft University in the Netherlands. They showed a video of a grad student from Russia doing a 6 month research internship in Paris (developing a control algorithm for the inverted pendulum problem). Then the speakers and the audience talked about various types of symbolism present in the film.

I also went to two industry/government panel discussions: "Challenges Confronting Mechanical Systems with Emphasis on Intelligence," and "Industry & Government Perspective on Issues and Challenges for Robotics." James Albus gave a presentation at the first one called "Building Brains for Thinking Machines," which was a brief summary of his work on hierarchical control systems. I talked with Dr. Albus briefly before his talk about some of his books (Intelligent Systems and Engineering of Mind).

Saturday, September 02, 2006

Now Listening to...

...audio lectures from another Gerald Schneider course, "Animal Behavior," on MIT OpenCourseWare.

I'm really enjoying taking courses on my own schedule. I just fit an entire course (Neuroscience and Behavior) into two weeks. (Actually I only listened to 16 out of the 30 audio lectures available, but they covered the topics I cared about.) It's great because I usually get bored with normal courses by the end of the semester. I have a pretty good feel for when I'm in a learning mood, and it usually doesn't coincide with scheduled lecture times. Now I can just keep my iPod with me all the time and listen to a chunk of a lecture while I'm riding the bus, exercising, or waiting for software to compile. I don't think I could go back to the old way. I'm done with old school school.

Thursday, August 24, 2006

Are Dreams Undirected Planning Sequences?

I had a thought. Say we have mental models of the world available for planning, and we can direct our planning sequences (i.e. trains of thought) through imagined spaces in order to achieve some goal. That's in the awake state. During dreaming, there is reduced activity in the prefrontal cortex, an area that is usually associated with planning/judgment. Maybe when we dream, our minds move through the same imagined spaces, but they do so in an undirected manner because the high-level planning control center is turned off. It's like the car is still moving, but no one's in the driver's seat. Without the prefrontal cortex, we can't really judge the absurdity of different situations, which explains why dreams always seem normal until we wake up.

Currently Listening To...

...audio lectures from Gerald Schneider's course, "Neuroscience and Behavior," on MIT OpenCourseWare.

Main Areas of Investigation

The main areas of the Verve agent design that need work are:

  • Context representation - should be more computationally efficient, and it should include a short-term memory of previous observations
  • Hierarchical action construction - complex actions should be built up automatically from low-level primitives
  • Experimentation with curiosity - this will be much more interesting once the previous two areas have been developed

Tuesday, August 22, 2006

I'm back from New York

I just got back from my internship in New York on Sunday night. It was a lot of fun. I had some great discussions with the other researchers in my group. My project was to help develop and test a computational model of the cerebellum. I learned a lot about the cerebellum (of course) and was introduced to several information theory-based approaches (like Linsker's Infomax framework and the subsequent ICA work of Bell & Sejnowski).

Now that I'm home, I'll be able to work on my own research again. I'd like to start looking at better context representations soon, possibly using ICA.

Sunday, August 06, 2006

The Classical and Romantic Qualities of Jazz Improvisation

I saw Arturo Sandoval play at the Blue Note on Friday. It was great, of course. I first heard him play in high school when I bought his album Hot House.

As I was listening, I started thinking about what goes on in your head when you improvise jazz. I think it helps to separate things into two categories: playing well technically (high notes, fast difficult patterns, etc.), and playing well emotionally (having a good gut feeling for what sounds good when). In other words, knowing how to play and knowing what to play. These two components trade control back and forth throughout a solo; one part chooses which pattern to play next, the other part executes it, the first part chooses another pattern, the other part executes it...

Some improvisers are good at one but not the other. For example, some are very proficient at playing difficult material, like high notes and fast, intricate syncopated sequences. But they don't have any clear direction to their solos (i.e. they're not good at choosing patterns). Others lack technical skill (i.e. they haven't mastered a lot of patterns), but they choose sounds in a way that produces an interesting overall message. They play with soul. I like to think of these two types of skill as Classical and Romantic Quality, as defined by Robert Pirsig in Zen and the Art of Motorcycle Maintenance. (I also think these two categories can be mapped roughly onto the brain's cognitive/analytical processes and its emotional/limbic processes, respectively.)

The best imrovisers are good at both kinds of Quality. Their brains have large libraries filled with high-quality patterns (Classical), and they have amazing selection mechanisms (Romantic). Their solos have relevant messages, and they're constructed out of masterful building blocks.

I haven't played jazz for a little while (about a year and a half), but I'd like to play again. I stopped playing because grad school got busy, mainly, but I also became frustrated with how I was playing. One thing was that I didn't like how I sounded in recordings. Something just wasn't right. Kind of like how most people don't like to hear their own voice. Another thing was that I wasn't really advancing as an improviser. I think it was because I was too focused on achieving Classical Quality. I would always practice with the goal of mastering some new pattern in all twelve keys, for example. I was never able to focus on learning to fashion the overall structure of a solo into anything meaningful. Then when it came time to play a solo with a group, I would go on autopilot, randomly jumping from one well-learned pattern to the next. My solos lacked direction.

I think what led to this problem was how, in most school bands I played in, the kids that could play the highest/fastest solos got the most respect (usually from their peers, not necessarily from the directors). So people that wanted to be noticed (myself included) would practice playing high notes and flashy passages. The Classical Quality of musical performance alone was usually equated with proficiency. Of course, it's hard to practice (and harder to teach) the Romantic Quality of jazz improvisation, especially in isolation. It's best to get real experience with a group that's not totally focused on the attention-getting parts of soloing.

Pirsig's division of idealogies into Classical and Romantic appears to be a good one. It appears to encompass a wide range to situations, including jazz improvisation. (I think that this division can be tied somewhat directly to the structure and function of our brains. I'm interested to find out if this is true.) Some improvisers choose to focus on developing one are more than the other (sometimes to the point of producing totally unbalanced solos), but the best soloists are good at both.

Tuesday, August 01, 2006

Curiosity Rewards Related to Flow

First, here's a diagram I made for a presentation on Artificial Curiosity at the HCI Forum 2006. It shows the simple idea behind the powerful Schmidhuber type of curiosity mechanism:


I had read about Mihaly Csikszentmihalyi's idea of "flow" in the book Satisfaction: The Science of Finding True Fulfillment, by Gregory Berns (which totally sounds like a hokey self-help book, but it's not). According to wikipedia, "The flow state is an optimal state of intrinsic motivation, where the person is fully immersed in what he or she is doing." Yesterday I came across Jenova Chen's game flOw, which is supposed to be based on the same principle. He has this great diagram in his thesis that explains flow quite simply:


Notice any similarity between this and Schmidhuber's curiosity model? Seems like the same type of mechanism; the optimal situation is somewhere between overly challenging and too simple. Intrinsic rewards for situations with maximum learning progress. Or, in the case of flow, intrinsic rewards for appropriate challenges. This seems like a major isomorphism.

Monday, July 31, 2006

Simulacra and Simulation Related to Online Virtual Worlds

From Simulacra and Simulation, by Jean Baudrillard, 1981, page 1:
"If once we were able to view the Borges fable in which the cartographers of the Empire draw up a map so detailed that it ends up covering the territory exactly (the decline of the Empire witnesses the fraying of this map, little by little, and its fall into ruins, though some shreds are still discernible in the deserts - the metaphysical beauty of this ruined abstraction testifying to a pride equal to the Empire and rotting like a carcass, returning to the substance of the soil, a bit as the double ends by being confused with the real through aging) - as the most beautiful allegory of simulation..."
This particular idea of simulation, a copy of some original thing which might acquire more attention than the original itself, makes me think of what Google Earth might turn into.
"...this fable has now come full circle for us, and possesses nothing but the discrete charm of second-order simulacra.

Today abstraction is no longer that of the map, the double, the mirror, or the concept. Simulation is no longer that of a territory, a referential being, or a substance. It is the generation by models of a real without origin or reality: a hyperreal. The territory no longer precedes the map, nor does it survive it. It is nevertheless the map that precedes the territory - precession of simulacra - that engenders the territory, and if one must return to the fable, today it is the territory whose shreds slowly rot across the extent of the map. It is the real, and not the map, whose vestiges persist here and there in the deserts that are no longer those of the Empire, but ours. The desert of the real itself."
The simulacrum concept (a copy with no original), however, appears to be more similar to our fantasy virtual worlds: Second Life, World of Warcraft, There, Project Entropia.

The major distinction between these two types of online worlds (that one is based on our physical world and the other is totally fictional) might become more pronounced in the future as these places gain in popularity. Each type has its benefits: one is instantly familiar, the other allows more creative freedom. Both seem to have their own place in the foreseeable future. If people someday spend so much time in virtual worlds that the real one is no longer familiar, the Earth-simulation type might fade away.

Thursday, July 13, 2006

Ethical Issues in Advanced Artificial Intelligence

Why should we study intelligence (either artificial or biological) and develop intelligent machines? What is the benefit of doing so? About a year ago I decided my answer to this question:

To the extent that scientific knowledge is intrinsically interesting to us, the knowledge of how our brains work is probably the most interesting topic to study. To the extent that technology is useful to us, intelligent machines are probably the most useful technology to develop.

I have been meaning to write up a more detailed justification of this answer, but now I don't have to, because I just read this great paper... "Ethical Issues in Advanced Artificial Intelligence," by Nick Bostrom. I think this paper provides a clear explanation for developing intelligent machines. It touches on all the important issues, in my opinion.

A few points really resonated with me. Here are some succinct excerpts:
  1. "Superintelligence may be the last invention humans ever need to make."
  2. "Artificial intellects need not have humanlike motives."
  3. "If a superintelligence starts out with a friendly top goal, however, then it can be relied on to stay friendly, or at least not to deliberately rid itself of its friendliness."
The third point, I believe, adequately deals with the issue of intelligent machines going crazy and taking over the world. It's all about motivation. If a machine has no intrinsic motivation to harm anybody, it will not do so. There are some caveats to this, some of which were discussed in the paper. I don't think any of these are insurmountable. First, random exploration during development in an unpredictable world will inevitably cause damage to someone/something. I don't think this is a huge problem, as long as the machine is sufficiently contained during development. Second, a machine with a curiosity-driven motivation system will essentially be able to create arbitrary value systems over time. The solution to this is simply to scale the magnitude of any "curiosity rewards" to be smaller than a hard-wired reward for avoiding harm. Third, a machine that can change its own code might change its motivations into harmful ones. Again, hard-coding a pain signal for any code modifications would help combat the problem. Further, if any critical code or hardware is modified, the whole machine could shut itself down. Of course, malignant or careless programmers might build intelligent machines with harmful motivations, but this is a separate issue from statement #3 above.

Tuesday, July 11, 2006

July Update

I've been pretty busy with my internship this summer, so I haven't done much work with Verve. It's probably best that I wait till I'm done at IBM, anyway, to avoid any conflicts of interest. The internship itself is going well. I'll post more details when it's over... maybe after we've submitted some things to be published.

Next week is the AAAI 2006 conference in Boston. I'll be there to present a poster in the student abstract program. My abstract is entitled, "Curiosity-Driven Exploration with Planning Trajectories." More details are here.

Wednesday, June 21, 2006

Some Biological Evidence of Curiosity Rewards

I can't stop thinking about curiosity rewards and how interesting the whole concept is. It seems like such a fundamental component of mental development. I can't wait to implement the more powerful model in Verve and run some experiments.

I just came across this article (referring to an article in the latest issue of American Scientist, which I couldn't find) describing some recent biological evidence of a sort of curiosity reward system: http://www.physorg.com/news70030587.html

The original (not free) article is in American Scientist... here is the first paragraph and a good diagram: http://www.americanscientist.org/template/AssetDetail/assetid/50770

Evidently, there is a "25-year-old finding that mu-opioid receptors – binding sites for natural opiates – increase in density along the ventral visual pathway..." These receptors are said to be involved in pain/pleasure processing elsewhere. Based on this evidence, Irving Biederman ran experiments using fMRI on humans viewing various types of images. "Biederman also found that repeated viewing of an attractive image lessened both the rating of pleasure and the activity in the opioid-rich areas."

It seems that the evidence described here might correspond to the more powerful model of curiosity, which says that curiosity rewards are proportional to the reduction in novelty. So, once the stimulus is fully predicted (and no longer novel), the rewards will disappear.

Here is another good quote: "The brain's craving for a fix motivates humans to maximize the rate at which they absorb knowledge."