Autonomous Systems & RLMay 2025 – Present (~16 months in production)Lead / Core Engineer · Autonomous QA

Roku Explorer

Autonomous Exploratory Testing Platform

A production autonomous QA platform for firmware and Grand Central product areas. Combines PPO/LSTM reinforcement learning, ChromaDB multimodal RAG, dual validators, and MCP-based device control to navigate device UIs, surface potential issues, and flag visual regressions with minimal human oversight.

249,868Device actions
1,065Runs
92Devices
30+Users
16,266Validation signals
1,993 / 5,966Nodes / transitions
The Problem & Engineering Constraint

The Core Challenge

Manual exploratory testing across complex streaming device interfaces is labor-intensive, error-prone, and struggles to scale across thousands of UI states, firmware variations, and rapid release cycles.
Technical Architecture & Approach

Engineering Solution & Implementation

Designed a multi-service Docker Compose stack (API, worker, UI, dual validators, scheduler, MCP) around an autonomous agent loop: PPO/LSTM policies plus multimodal vision reasoning explore device UIs, while change-driven planning targets high-risk paths and MCP commands physical test devices. UI states are indexed in ChromaDB for retrieval-augmented reasoning across firmware and Grand Central product areas.

SYSTEM ARCHITECTURE: ROKU EXPLORER
CHANGE-DRIVEN AUTONOMOUS QA PIPELINE
1. REPO DIFF• Git Pull Request Diffs• Code Surface Impact• Changed Component Map2. SCENARIO SYNTHESIS• LLM Code Understanding• At-Risk UI State Targets• Focused Test Spec Enqueue3. EXPLORER CORE ENGINERL Policy (PPO / LSTM)Action selection & exploration bonusChromaDB UI Vector MemoryScreen state similarity & historyMultimodal Vision ReasonerOCR, UI bounds, visual anomaly check4. DEVICE FLEET• Real Roku Hardware• Keystroke Ingestion• Live HDMI Frame GrabScreen Frames5. BUG & RISK TRIAGE• Functional Crash Log• Visual Regressions FlaggedDefect Telemetry

Figure 1: High-level architectural dataflow of Roku Explorer's change-driven autonomous loop.

Measured Production Impact

Verified Outcomes & Deliverables

In production for ~16 months with 249,868 device actions across 1,065 runs on 92 devices, serving 30+ users and surfacing 16,266 validation signals over a 1,993-node / 5,966-transition exploration graph.

Designed and still lead the platform — roughly 70% of commit history across a 17-engineer contributor base (~4,000 commits).

Closed the loop from code change → automated risk scenario generation → device exploration with dual validators and MCP device control.

Technologies & Components

System Tooling & Technologies

PythonPyTorch (PPO / LSTM)ChromaDBFlaskReact 18Docker ComposeMCPNabu CloudOpenCVLiteLLM