How we built our Reinforcement Learning Environment
Abstract. We built a Reinforcement Learning Environment (RLE) to explore BIOBUZZ strategy before committing ideas to a physical robot. The environment combines a CAD-informed 2D field, headless Python physics, legality constraints, multi-agent policy learning and recorded match playback. MAPPO searches for promising coordinated behavior; Monte Carlo analysis tests how that behavior performs across repeated simulated conditions. These are two different jobs: finding a candidate and deciding whether the evidence supports it. [1][3]
Our inspiration came from DeepMind's work learning Pong and Breakout through interaction, and from Mustafa Suleyman's hill-climbing mindset: take a step, measure it, keep what improves, and learn from what fails. We apply that mindset to both strategy and the simulator itself. This paper explains the build process, not a new score record, a proof of an optimal policy, or a calibrated prediction of robot performance. [1][2]
See what the environment records
An existing full-match study recording, seed 32261043, demonstrates the 2D replay system. Scripted Red reaches 13 TIPs and 271 points; the Blue training alliance reaches nine TIPs and 186 points. This is an illustrative retained witness, not a new experiment or evidence that MAPPO learned the best sequence. The recording includes passive settling after 150 powered seconds.
Loading the recorded match...
[*] Starting score: Each alliance begins with 10 modeled field-content points: 4 GARDEN points (four starting POLLEN at 1 point each) plus 6 HIVE-retained points (three starting NECTAR at 2 points each). These are already present at setup, not earned by robot actions. We keep the original recorded totals; contents points change as pieces move and are not a fixed 10-point bonus. Replay evidence.
Starts paused at 1x. Positions are original recorded frames, not interpolated or resimulated. Exact TIP event timestamps can fall between frames; event jumps show the first recorded frame at or after completion.
Playback provenance
Original recording information appears when playback loads.
Download the sanitized replay dataSource [4]. Uncalibrated simulation; no claim of physical robot performance.
Discover. Test. Inspect. Improve.
Model the game
Use field CAD to inform geometry, then represent movement, pieces, scoring and phase timing in a 2D physics environment.
Constrain interaction
Place physics and safety controllers between a robot's proposed behavior and the resulting game state. An impossible move is not a useful strategy.
Learn with MAPPO
Collect multi-robot experience and update policies using a centralized training coach with individually acting robots.
Evaluate with Monte Carlo
Test candidate behavior repeatedly under varied simulated conditions and opposition; separate a memorable success from a repeatable improvement.
Inspect and climb
Use retained actor logs and replays to explain outcomes, expose model flaws and guide the next measured change.
1. From Pong and Breakout to a strategy laboratory
DeepMind's Atari work showed why reinforcement learning is more than scripting a clever move. An agent observes a game, chooses an action, receives a reward and updates its behavior from experience. In Pong, repeated interaction can improve paddle decisions. In Breakout, learned play can uncover useful patterns such as opening a route behind the bricks. The exciting part for us was strategy emerging from feedback instead of every useful sequence being written by hand. [1][2]
We borrowed that experimental idea, not the exact Atari implementation. DeepMind's early work used deep Q-learning and screen-based input; our BIOBUZZ work uses a multi-agent PPO approach and a modeled field. Neither the algorithm nor the observation interface should be treated as identical. The common principle is to make a game cheap to repeat, let policies explore, and judge the result against evidence. [2][3]
Our question is practical: which decisions help an alliance use its time and game pieces well? Collection, travel, shooting, HIVE access and parking compete for the same resources. A simulation lets us investigate those tradeoffs without wearing out hardware on every exploratory attempt. It narrows the questions for physical tests; it does not replace them. [1]
2. Start with the field, then make it behave
We began with field CAD and used Blender in the geometry and validation workflow. The strategy runtime became a headless, CAD-informed Python physics model with a 2D view for inspection. Geometry gives the environment a spatial foundation, but a recognizable field is not yet a trustworthy environment: pieces must move, collisions must matter and scoring must follow modeled state changes. [1]
The first prototype made that distinction obvious. It exposed idle behavior, incorrect field orientation, robots passing through structures, teleportation and missing GARDEN gathering. A high score in that version could reward an implementation mistake rather than a good decision. Correcting the world was therefore part of the learning project, not cosmetic work after training. [1]
The later release candidate described four equal-capability robots, 56 game pieces, 30 seconds of AUTO and 120 seconds of TELEOP actuation, with physical collisions, ball flight, HIVE tipping and parking represented. Those are configuration facts, not a claim of complete rule coverage. The official transition interval is not represented by those 150 powered seconds, and individual research studies can restrict the actions or objectives further. [1][4][5]
Two-dimensional strategy simulation also does not mean ignoring every airborne interaction. The model includes ball-flight behavior while replay presents the field from above. The development account does not provide enough equations or parameters to claim a particular integrator, collision algorithm or fully calibrated three-dimensional dynamics. [1]
3. Define the agent-environment contract
A learning environment needs a repeatable interaction loop: initialize a match, provide an observation, receive proposed actions, advance the model, compute feedback and identify when the episode ends. That is the conceptual contract we use to explain the system; it is not a published API specification. The source presentation does not enumerate the exact reset interface, tensor shapes, timestep or action encoding. [1]
The training account describes individually acting robots with limited local information and a coach that sees the full field during training. We distinguish an agent's observation from the trainer's broader state information. The presentation's local-vision explanation does not establish a particular camera, sensor model or deployment-ready perception stack. [1][3]
Moving, gathering, shooting, tipping and parking are the behaviors the environment seeks to represent. Physics and safety controllers sit beneath policy learning so that a proposed action cannot become a valid strategy simply by crossing a wall or teleporting. These modeled gates are not a complete competition referee or a physical safety certification. [1]
Rewards connect experience to the objective. The coach is described as evaluating how decisions help the team score, but the exact numerical reward function is not published in the presentation. Training reward, official scoring and evaluation criteria must not be conflated: a shaped reward is a learning signal, not automatically a competition point. We do not invent reward weights or claim that the current formulation eliminates reward exploitation. [1][3]
4. MAPPO searches for coordinated behavior
We use Multi-Agent Proximal Policy Optimization (MAPPO) to search for the best outcomes we can find under the modeled conditions. Multi-agent matters because one robot's choices change its partner's opportunities: collecting, reloading and approaching a HIVE are not independent scheduling problems. A policy that looks good alone may interfere with an alliance partner. [1][3]
MAPPO follows centralized training with decentralized execution. In the standard formulation, actors choose actions from their own observations while a centralized critic can use broader state information to estimate returns during training. The presentation calls that training role the coach. It should not be mistaken for a full-field controller secretly making every robot's decisions during execution. [1][3]
PPO-style updates use a clipped optimization objective to discourage overly large policy changes from one batch of experience. In plain language, learning should not abandon useful behavior because one unusual episode went well. Clipping is a stabilizing technique, not a guarantee that every update improves the policy. Network architecture, learning rates, clipping settings and training budgets are not specified here. [3]
The best outcome observed in a search is not necessarily an outcome learned by MAPPO. Our published full-match study makes this visible: all seven safety-qualified 13-TIP alliance outcomes came from scripted opponents, while the learning alliance reached at most 12. That distinction is essential. The environment demonstrated a modeled possibility; it did not demonstrate a learned policy that reliably achieves it. [4]
5. Monte Carlo tests the candidate, not the headline
After strategy discovery, we use Monte Carlo analysis to test candidate outcomes across repeated simulated conditions. The presentation describes a tester that varies lucky and unlucky factors and defensive scenarios, a challenger that supplies opposition, and a grader that evaluates unseen matches before promotion. MAPPO changes the policy from experience; Monte Carlo estimates how a chosen policy behaves under a specified sampling setup. Monte Carlo is not itself the policy-learning algorithm. [1]
One spectacular match is a witness of possibility. A distribution of matches answers a different question: how often does the behavior work, how variable is the result, and where does it fail? Mixing episodes from continually changing training policies can describe a search history, but it cannot by itself estimate the reliability of a single frozen policy. Our existing studies explicitly separate exploratory outcomes from frozen promotion evaluations. [4][6]
A stronger evaluation protocol freezes the candidate and incumbent, uses held-out seeds, compares them on matched conditions and balances alliance color or starting position where relevant. Useful reporting includes the full denominator, mean and spread, success rate for a stated target, and the frequency of geometry or safety failures. These are methodological recommendations, not claims that every development run used the same complete protocol. [6]
Safety filtering needs equal care. Report rejected matches as well as the retained population; otherwise a policy that often fails the constraints can look strong when only its surviving games are shown. If many candidates are selected on the same test set, selection can overfit that set, so a final held-out evaluation should remain separate. More samples reduce sampling uncertainty, but do not remove bias from an unrealistic environment. This methods paper reports no new confidence interval or new Monte Carlo run. [4][6]
6. Make experiments fast and outcomes inspectable
Throughput determines how many ideas we can afford to test. The development account describes a pivot toward headless physics and multiple worker processes rather than relying on Python threads for CPU-heavy collection. GPU/DirectML and AMD NPU inference were explored, but physics remained an important CPU bottleneck. Accelerator access alone does not make the whole simulation fast; the slowest substantial part of the loop still matters. [1]
We also changed what a replay stores. Early runs produced large MP4 recordings. Later work retained structured digital actor logs that could drive HTML Canvas playback. This separates match simulation from rendering: we can inspect recorded robot positions, game-piece motion and completion events without rerunning the match or treating a video as the only evidence. [1][4]
The embedded recording is an example of that evidence workflow. Playback starts paused, supports frame stepping and event jumps, and uses retained coordinates rather than inventing a new trajectory. It preserves the original state-based score, including 10 modeled field-content points per alliance at setup; those contents are not robot-earned actions or a fixed bonus. The linked study explains the source hashes and score interpretation. [4]
Selective recording has a tradeoff. Keeping detailed logs only for qualifying high-scoring matches makes the evidence archive smaller, but those highlights are not an unbiased sample. Retained witnesses explain a sequence; complete statistics and separately designed evaluations are needed to discuss reliability. Missing recordings should remain missing, not be reconstructed and presented as observed motion. [1][4][6]
7. Hill climbing improved the laboratory as well as the policy
Mustafa Suleyman's hill-climbing mindset is our inspiration for disciplined iteration, not a claim that he designed or endorsed this environment. Our working interpretation is simple: change something, measure what changed, retain the improvement and revisit failures. This is a development mindset; MAPPO's gradient-based policy optimization and the external promotion loop remain distinct mechanisms.
The development presentation traces nine iterations through CAD and Blender foundations, an early strategy prototype, mutation-based search, physics corrections, MAPPO, throughput experiments, replay redesign, score-bound studies and a source-pinned study series. The progression was not just toward larger numbers. It was toward stronger reasons to trust or reject a number. [1]
That gives the project two coupled loops. The inner loop learns policies inside a version of the environment. The outer loop examines physics, constraints, evaluation and evidence, then decides what to improve or promote. A better-looking policy score may be caused by a model change, so versions and assumptions must travel with the result. Preserved studies should remain tied to their original evidence rather than silently acquiring a new interpretation. [1][4][6]
8. What the RLE gives us, and what comes next
The RLE gives Cookie Coders a repeatable place to ask strategy questions and a way to inspect why a sequence succeeded. It can reveal coordination opportunities, expose simulator defects and prioritize physical experiments. It cannot turn a modeled score into a robot record, establish a global optimum, or remove the need to measure hardware. [1][4]
Important assumptions remain. The presentation describes collection in all directions and scoring from positions with line of sight; the published studies also name idealized sensors, equal hardware and uncalibrated shot and HIVE-motion parameters. Those assumptions can make a strategy easier in simulation than on a real robot. Calibration needs measured movement, collection, shooting, HIVE dynamics and representative opposition before transfer claims become credible. [1][4][6]
A reproducible implementation paper should additionally pin code and dependency versions, observation and action definitions, the numerical reward function, physics parameters, MAPPO configuration, seeds, evaluation splits and promotion criteria. The presentation supplies a development account, not that full reproduction package. Students own the next questions and explanations; physical robot experiments remain adult-supervised. Our conclusion is deliberately modest: learn in simulation, test across variation, inspect the evidence, then measure on the robot.
The build, separated by responsibility
A conceptual architecture summary grounded in the development presentation, not a benchmark or an executable interface specification.
| Layer | Responsibility | What it does not establish |
|---|---|---|
| CAD and geometry | Inform the modeled field and spatial constraints. | As-built robot or field calibration. |
| Physics and safety controllers | Advance interactions and reject modeled impossible behavior. | Complete rule compliance or physical safety certification. |
| MAPPO and coach | Learn candidate multi-robot policies from experience. | A global optimum or guaranteed improvement. |
| Tester, challenger and grader | Evaluate variation, opposition and promotion evidence. | Real-world reliability outside the sampled model. |
| Actor logs and replay | Make retained trajectories and events inspectable. | An unbiased performance distribution from highlights alone. |
Source [1].
What this paper does not prove
- This is a methods and development account. It does not add a training run, a Monte Carlo benchmark, a score record or a proven maximum.
- The supplied presentation does not publish exact observation tensors, action encoding, reward weights, network architecture, physics equations or hyperparameters.
- The retained demonstration is from an existing study; the 13-TIP sequence belongs to a scripted opponent, not a demonstrated MAPPO champion.
- Simulation assumptions and safety filters remain uncalibrated and incomplete. Statistical confidence inside a model is not confidence in physical transfer.
- The Atari research and hill-climbing mindset are inspirations, not identical implementations or endorsements. BloxBuzz engine, rules and physics parity remains a goal.
Where to play BIOBUZZ
BloxBuzz is where you can play BIOBUZZ; the RLE is where our agents train and are evaluated. The RLE does not run in Roblox, and our goal is for it to use the same game engine, rules and physics as BloxBuzz. These results come from the RLE, not from BloxBuzz gameplay, and this study does not verify that the two match.
Opens Roblox in a new tab. Roblox's age requirements and privacy practices apply. BloxBuzz is unofficial and is not affiliated with or endorsed by FIRST or RTX.
Play BloxBuzz on RobloxSources and evidence
This methods account paraphrases the supplied October 9 development presentation and the team's stated inspirations. The presentation remains privately retained; its slides 7 and 9-15 ground the development narrative. Public papers explain the algorithms, while the linked Cookie Coders studies supply the retained example and evaluation distinctions. We do not publish the private deck or invent a new evidence JSON for this account.
- Development presentation: environment, learning and evaluation
Privately retained project presentation, supplied October 9, 2026. Slide 7: RL inspiration; slides 9-10: prototype and release candidate; slides 12-14: MAPPO, evaluation roles and Monte Carlo; slide 15 and speaker notes: iteration history, physics, throughput and replay. This in-page link leads to the development account, not a public copy of the deck.
- Mnih et al. (2013), Playing Atari with Deep Reinforcement Learning
Primary DeepMind deep Q-learning paper. Background for trial-and-error Atari learning; not evidence that BIOBUZZ uses DQN, pixel observations or the same training budget.
- Yu et al. (2022), The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games
Primary MAPPO reference, first posted in 2021. Explains cooperative multi-agent PPO and centralized training; not documentation of this project's exact network or hyperparameters.
- Cookie Coders full-match study and retained replay evidence
Existing 50,000-match study, scripted-versus-learning attribution, safety-filter denominator, state-based setup score and source-hashed recording for seed 32261043. Reused as a demonstration, not a new result.
- Official BIOBUZZ Competition Manual
Consult the current official revision before physical tests. Modeled 30-second AUTO and 120-second TELEOP actuation do not reproduce every timing or refereeing detail.
- Cookie Coders solo AUTO study: search versus frozen evaluation
Existing source-pinned study separates changing exploratory actors from generation-19 frozen promotion evaluation, including rejected candidates and retained-evidence limits.