

{"id":150,"date":"2015-08-26T18:10:27","date_gmt":"2015-08-26T16:10:27","guid":{"rendered":"https:\/\/project.inria.fr\/ExTra-Learn\/?p=150"},"modified":"2016-05-27T13:59:27","modified_gmt":"2016-05-27T11:59:27","slug":"two-papers-on-apprenticeship-learning-at-ijcai15","status":"publish","type":"post","link":"https:\/\/project.inria.fr\/ExTra-Learn\/two-papers-on-apprenticeship-learning-at-ijcai15\/","title":{"rendered":"Two papers on apprenticeship learning at IJCAI&#8217;15"},"content":{"rendered":"<p>We presented two novel results on apprenticeship learning at IJCAI&#8217;15, where we\u00a0use demonstrations generated by an expert to learn near-optimal policies. This scenario is typical in source-to-target transfer, where the expert policy could be the optimal policy in a source task and we want to learn a policy in the target task.<\/p>\n<p class=\"p1\"><span class=\"s1\"><b>Direct Policy Iteration with Demonstrations <\/b>(Jessica Chemali, Alessandro Lazaric) [<a href=\"https:\/\/project.inria.fr\/ExTra-Learn\/files\/2016\/05\/DPID_CameraReady.pdf\">pdf<\/a>]<\/span><\/p>\n<p class=\"p1\"><span class=\"s1\"><i>We consider the problem of learning the optimal policy of an unknown Markov decision process (MDP) when expert demonstrations are available along with interaction samples. We build on classification-based policy iteration to perform a seamless integration of interaction and expert data, thus obtaining an algorithm which can benefit from both sources of information at the same time. Furthermore, we provide a full theoretical analysis of the performance across iterations providing insights on how the algorithm works. Finally, we report an empirical evaluation of the algorithm and a comparison with the state-of-the-art algorithms.<\/i><\/span><\/p>\n<p class=\"p1\"><span class=\"s1\"><b>Maximum Entropy Semi-Supervised Inverse Reinforcement Learning <\/b>(J. Audiffren, M. Valko, A. Lazaric, and M. Ghavamzadeh) [<a href=\"https:\/\/project.inria.fr\/ExTra-Learn\/files\/2016\/05\/messi-TR.pdf\">pdf<\/a>]<\/span><\/p>\n<p class=\"p1\"><span class=\"s1\"><i>A popular approach to apprenticeship learning (AL) is to formulate it as an inverse reinforcement learning (IRL) problem. The MaxEnt-IRL algorithm successfully integrates the maximum entropy principle into IRL and unlike its predecessors, it resolves the ambiguity arising from the fact that a possibly large number of policies could match the expert&#8217;s behavior. In this paper, we study an AL setting in which in addition to the expert&#8217;s trajectories, a number of unsupervised trajectories is available. We introduce MESSI, a novel algorithm that combines MaxEnt-IRL with principles coming from semi-supervised learning. In particular, MESSI integrates the unsupervised data into the MaxEnt-IRL framework using a pairwise penalty on trajectories. Empirical results in a highway driving and grid-world<span class=\"Apple-converted-space\">\u00a0 <\/span>problems indicate that MESSI is able to take advantage of the unsupervised trajectories and improve the performance of MaxEnt-IRL. <\/i><\/span><\/p>\n<p><\/p>","protected":false},"excerpt":{"rendered":"<p>We presented two novel results on apprenticeship learning at IJCAI&#8217;15, where we\u00a0use demonstrations generated by an expert to learn near-optimal policies. This scenario is typical in source-to-target transfer, where the \u2026<\/p>\n<p class=\"continue-reading-button\"> <a class=\"continue-reading-link\" href=\"https:\/\/project.inria.fr\/ExTra-Learn\/two-papers-on-apprenticeship-learning-at-ijcai15\/\">Continue reading<i class=\"crycon-right-dir\"><\/i><\/a><\/p>\n","protected":false},"author":552,"featured_media":151,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_members_access_role":[],"_members_access_error":""},"categories":[9],"tags":[],"class_list":["post-150","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-publications"],"_links":{"self":[{"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/posts\/150","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/users\/552"}],"replies":[{"embeddable":true,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/comments?post=150"}],"version-history":[{"count":3,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/posts\/150\/revisions"}],"predecessor-version":[{"id":170,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/posts\/150\/revisions\/170"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/media\/151"}],"wp:attachment":[{"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/media?parent=150"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/categories?post=150"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/project.inria.fr\/ExTra-Learn\/wp-json\/wp\/v2\/tags?post=150"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}