Gokul Swamy's Avatar

Gokul Swamy

@gokul.dev

PhD student at @cmurobotics.bsky.social working on efficient algorithms for interactive learning (e.g. imitation / RL / RLHF). no model is an island. prefers email. https://gokul.dev/. on the job market!

4,302
Followers
417
Following
109
Posts
11.11.2024
Joined
Posts Following

Latest posts by Gokul Swamy @gokul.dev

Couldn't think of a better excuse than #neurips2025 to visit my home town of San Diego! I'll be around 12/3-12/6 -- shoot me a DM / email if you'd like to catch up!

26.11.2025 17:12 πŸ‘ 1 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0
Preview
The Game Theory of How Algorithms Can Drive Up Prices | Quanta Magazine Recent findings reveal that even simple pricing algorithms can make things more expensive.

Our paper on algorithmic collusion was featured in a Quanta article! www.quantamagazine.org/the-game-the...

22.10.2025 15:19 πŸ‘ 29 πŸ” 9 πŸ’¬ 2 πŸ“Œ 4

Woooaaaah :O

22.10.2025 17:51 πŸ‘ 1 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0
β€œRaising Our Sights (A Long Rant From an Accidental Engineer)”, Scott Shenker
β€œRaising Our Sights (A Long Rant From an Accidental Engineer)”, Scott Shenker YouTube video by Plamadiso - Platforms, Markets & Digital Society

Just discovered this lovely talk from my favorite professor from undergrad, who is still as inspiring as I remember him:
www.youtube.com/watch?v=XLZ0...

22.10.2025 02:30 πŸ‘ 0 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0

Please apply or help out!

13.10.2025 00:36 πŸ‘ 1 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0
Preview
The Epic Story of Maximum Likelihood At a superficial level, the idea of maximum likelihood must be prehistoric: early hunters and gatherers may not have used the words ``method of maximum likelihood'' to describe their choice of where a...

Late, but arxiv.org/abs/0804.2996 is *incredible*, so many good lines (e.g., "This comes close to being an accusation of a false claim of priority for a false discovery of an untrue fact, which would be a rare triple-negative in the history of intellectual property disputes.").

28.09.2025 21:13 πŸ‘ 12 πŸ” 3 πŸ’¬ 0 πŸ“Œ 0

I've been really enjoying the new Ninajirachi album -- it's very Boiler Room-core :)

23.08.2025 18:19 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

Thanks for the shout-out and I hope the lectures were at least somewhat understandable! Yeah, once things settle down a bit for me, I'd like to more deeply understand the connection between Rust's structural estimation and IRL as I conceive of it.

23.08.2025 18:16 πŸ‘ 2 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

We therefore advocate for caution when making or evaluating claims about LLM reasoning and beyond with GRPO and PPO, ideally using algorithms like RLoo or REBEL instead. Check out our blog post for links to our code and W&B logs if you'd like to reproduce our experiments.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0

While this worked out for the better on some seeds, it doesn't have to in general. After all, an algorithm that behaves unexpectedly *well* in one setting can perform unexpectedly *poorly* in another, perhaps more important, setting.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

We see similar results on a didactic bandit problem -- i.e. a problem that has nothing to do with LLMs or reasoning! This implies that PPO / GRPO are fundamentally *not* following the true policy gradient.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

We find that RLoo (an unbiased estimate of the vanilla PG) and REBEL (a regression-based approximation of online mirror descent) preserve performance as expected. In contrast, algorithms like PPO / GRPO that include heuristics (e.g. clipping) show a marked and unexpected change in performance.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

So, with a truly random reward function, all policies look equally good. Thus, the *true* policy gradient is zero, as the initial policy is optimal by construction. So, we'd expect performance to flatline. We use random rewards as a *diagnostic task* to compare different RL algs.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

Lead by Owen Oertell & Wenhao Zhan, joint w/ Steven Wu, Kiante Brantley, Jason Lee, and Wen Sun. If a project has got Wen, Owen, Wenhao, and Qwen on it, you know it's gotta be good πŸ˜›.

15.07.2025 17:46 πŸ‘ 1 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0
Preview
Heuristics Considered Harmful: RL With Random Rewards Should Not Make LLMs Reason | Notion Owen Oertell*, Wenhao Zhao*, Gokul Swamy, Zhiwei Steven Wu, Kiante Brantley, Jason Lee, Wen Sun

Recent work has seemed somewhat magical: how can RL with *random* rewards make LLMs reason? We pull back the curtain on these claims and find out this unexpected behavior hinges on the inclusion of certain *heuristics* in the RL algorithm. Our blog post: tinyurl.com/heuristics-c...

15.07.2025 17:46 πŸ‘ 6 πŸ” 2 πŸ’¬ 1 πŸ“Œ 0

very nice lectures, watch them from time to time

20.06.2025 06:07 πŸ‘ 6 πŸ” 1 πŸ’¬ 0 πŸ“Œ 0

Want to learn about online learning, game solving, RL, imitation learning with applications to robotics, and RLHF with applications to language modeling? Check out this course! πŸ‘

20.06.2025 13:11 πŸ‘ 6 πŸ” 1 πŸ’¬ 0 πŸ“Œ 0

While I can't promise everything will be crystal-clear after going though the lectures (especially because of my handwriting :p), I hope that if nothing else, you can tell how beautiful we all find these ideas. If that feeling comes across, I'll feel like I have succeeded! :)

20.06.2025 03:53 πŸ‘ 3 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

The second was being able to teach this course with my amazing advisors, Drew Bagnell and Steven Wu -- the folks I learned all of this stuff from. Fun fact: because of parking fees, Drew actually *paid* to lecture. And I'm always grateful to ZSW for pushing me out of the nest.

20.06.2025 03:53 πŸ‘ 3 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

Two other things made this course particularly special. The first was the students and their *incredible* questions -- there were so many times where I was like wow, it took me *YEARS* before I realized that was the right question to be asking.

20.06.2025 03:53 πŸ‘ 4 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0
Algorithmic Foundations of Interactive Learning SP25: Lecture 17
Algorithmic Foundations of Interactive Learning SP25: Lecture 17 YouTube video by Gokul Swamy

We also had wonderful guest lectures from Yuda Song
on hybrid RL (youtu.be/1B2XGXQ2hfA), Sanjiban Choudhury on scaling imitation (youtu.be/KnXSeTuCgFI), and Wen Sun on RLHF algorithms (youtu.be/qdkBZJywi_4).

20.06.2025 03:53 πŸ‘ 6 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0
Algorithmic Foundations of Interactive Learning SP25: Lecture 19
Algorithmic Foundations of Interactive Learning SP25: Lecture 19 YouTube video by Gokul Swamy

My favorite lectures to give were on the value of interaction in imitation / RLHF! youtu.be/uESAXg-CXFs, youtu.be/N8-Nh_iTmps, youtu.be/qHvB30J5gyo, youtu.be/ZzFjoH47GIg. It took 5 years, but I finally have an answer at least I find compelling :p.

20.06.2025 03:53 πŸ‘ 6 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

To do so, we worked backwards from things like ChatGPT and RMA and "backed out" a "dependency graph". We then did a "forward pass" over the semester, going from online learning, to game solving, to core RL, to imitation learning / robot learning, to RLHF / LLM fine-tuning.

20.06.2025 03:53 πŸ‘ 5 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0

I think in a field as fast-paced as machine learning, a good course gives students a conceptual framework for understanding new developments quickly + what is actually "new" vs. the classical algorithms. We also wanted to explain *when* scale isn't "all you need."

20.06.2025 03:53 πŸ‘ 6 πŸ” 0 πŸ’¬ 1 πŸ“Œ 0
Home Website for AFIL course.

You can access all the content here:
Course Website: interactive-learning-algos.github.io
Lecture Playlist: youtube.com/playlist?lis...
Scribe Notes "Book": interactive-learning-algos.github.io/assets/pdfs/....
Homeworks / class competition material are also public!

20.06.2025 03:53 πŸ‘ 14 πŸ” 1 πŸ’¬ 1 πŸ“Œ 0
Post image

It was a dream come true to teach the course I wish existed at the start of my PhD. We built up the algorithmic foundations of modern-day RL, imitation learning, and RLHF, going deeper than the usual "grab bag of tricks". All 25 lectures + 150 pages of notes are now public!

20.06.2025 03:53 πŸ‘ 47 πŸ” 10 πŸ’¬ 3 πŸ“Œ 2

Shortcut models enable scaling offline RL, both at train-time at test-time! We beat so many other algorithms on so many tasks we had to stick most of the results in the appendix πŸ˜…. Very proud of @nico-espinosa-dice.bsky.social for spearheading this project, check out his thread!

12.06.2025 23:14 πŸ‘ 1 πŸ” 1 πŸ’¬ 0 πŸ“Œ 0

Boston friends: I'll be in the Cambridge area for the next few days, shoot me a message if you'd like to catch up :).

27.04.2025 21:28 πŸ‘ 2 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0
Post image

I won't be at #ICLR2025 myself this time around but please go talk to lead authors Nico, Zhaolin, and Runzhe about their bleeding-edge algorithms for imitation learning and RLHF!

22.04.2025 14:05 πŸ‘ 4 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0
Preview
Efficient Imitation under Misspecification We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere. This is often true in practice due to ...

As always, I'm incredibly grateful to Wen / Sanjiban for letting me borrow their excellent students for a bit to work on my harebrained schemes. Full paper at arxiv.org/abs/2503.13162. [17/n, n=17]

07.04.2025 19:03 πŸ‘ 1 πŸ” 0 πŸ’¬ 0 πŸ“Œ 0