Under New York City’s Local Law 33, commercial properties must publicly display an annual Energy Grade Certificate—an A-to-F rating reflecting energy efficiency. For a leading NYC energy services provider managing thousands of properties, retrieving these compliance documents manually created a severe operational bottleneck.
To solve this challenge, we engineered an asynchronous, cloud-native browser automation platform on AWS. The architecture converted a multi-day manual effort into a resilient, cost-efficient pipeline that completes in hours—dynamically scaling worker capacity on demand and releasing compute resources as soon as processing finishes.
What’s in this article:
- Transforming a local Playwright script into a distributed, production-ready platform on AWS
- Overcoming anti-bot measures, dynamic site layouts, and silent browser crashes at scale
- Key architectural choices for high-throughput batch processing that keep AWS costs under control
The Operational Challenge
Retrieving an Energy Grade Certificate meant navigating a tedious, fully manual sequence on the NYC Department of Buildings (DOB) NOW Public Portal:
- Search: Query the portal using a property’s unique Borough-Block-Lot (BBL) identifier or Building Identification Number (BIN).
- Retrieve: Locate the record, generate the official Energy Grade Certificate PDF within the browser viewport, and trigger the download.
- Process & Store: Adjust document properties, apply standardized client naming conventions, and manually upload the final asset to centralized cloud storage.
At 2 to 3 minutes per document, the process was manageable for an isolated building. However, complex properties often house more than 20 distinct Building Identification Numbers (BINs), turning a portfolio of 1,500+ properties into over 100 hours of manual labor per reporting cycle.
Worse, the manual workflow was deeply fragile. A single connection drop, browser session timeout, or portal glitch wiped out active progress, forcing operators to restart from scratch with zero-error recovery and no capability to scale.
The Solution
Our goal was to eliminate the manual overhead of generating Energy Grade Certificates without introducing unnecessary system complexity or fragmented third-party tooling.
We built the platform natively on AWS to align seamlessly with the client’s existing environment. Leveraging core services like Amazon ECS Fargate, Simple Queue Service (SQS), Simple Storage Service (S3), and CloudWatch allowed us to deliver an enterprise-grade automation engine without introducing a secondary cloud vendor.
Automating the Browser
We began by writing a headless Playwright script to mimic human user interactions on the NYC DOB NOW portal. In isolated local tests, the script performed flawlessly: opening the target URL, navigating the portal UI, querying property IDs, waiting for DOM elements to render, and capturing the resultant PDF stream.
However, executing a script on a developer’s local workstation is different from running thousands of concurrent browser sessions in production. As we transitioned from single properties to portfolio-wide batches, we encountered fundamental technical hurdles inherent to headless browser scraping:
- HTTP Timeout Constraints: Synchronous web requests would time out while waiting for slow responses.
- Resource Exhaustion: Headless Chrome instances consume heavy memory and CPU footprint per context; spinning up dozens of concurrent browser tabs on standard instances quickly triggered out-of-memory (OOM) crashes.
- Lack of Fault Tolerance: Network drops or unexpected popups caused the entire script to crash, leaving no trace of which properties succeeded and which failed.
- State Management & Visibility: Operations teams had no real-time transparency into batch progress or system status.
Asynchronous, Cloud-Native Architecture
We decoupled the application into an event-driven, asynchronous background processing system using Amazon SQS and AWS Fargate. This helped to overcome the limits of synchronous execution.
How the Distributed System Works:
- Task Queueing (Amazon SQS): Large batches are chunked into individual property tasks and pushed to an SQS queue, acting as a high-capacity digital buffer.
- Parallel Execution (AWS Fargate): A fleet of containerized Playwright workers (Docker + Playwright) automatically spins up, pulls tasks off the queue, and executes processing against the DOB portal in parallel.
- Real-Time Transparency: Submitting a batch generates a unique Batch ID. Workers write outputs to Amazon S3 and update PostgreSQL in real time, allowing operators to monitor live progress via a status dashboard.

Production Optimizations
Decoupling execution solved API bottlenecks, but operating at scale required addressing platform resilience and infrastructure costs.
Platform Resilience & Self-Healing
In a production environment, third-party portal downtime, network drops, and browser crashes are inevitable. Rather than treating these failures as exceptions, we engineered self-healing mechanisms directly into the worker runtime:
- State Tracking: Workers record real-time execution status, stack traces, and failure metrics in PostgreSQL.
- Automated Retries: Transient network drops automatically trigger exponential backoff retry logic.
- Granular Recovery: When persistent errors occur, the failed tasks are exported to an SQS Dead-Letter Queue (DLQ), allowing operators to re-run specific items without restarting the entire batch.
Instead of manually digging through entire runs to troubleshoot drops, operators can pinpoint exact root causes in seconds.
Dynamic Worker Scaling & Cost Optimization
To handle large certificate-generation requests efficiently, we implemented workload-based smart scaling. The platform dynamically adjusts its team of digital workers both horizontally and vertically based on batch size. Small requests use minimal resources, while massive batches instantly spawn multiple parallel workers. Once processing wraps up, the workers scale back to zero, eliminating ongoing server overhead when idle.
Business Impact & Results
- Faster Turnaround Time: Processing ~1,400 properties dropped from 2–3 days down to 3–4 hours using parallel Fargate workers.
- Cloud Infrastructure Savings: Serverless containers scale out on demand and shrink to zero when queues clear, eliminating idle server costs.
- Near-Zero Manual Intervention: Automated retries via SQS Dead-Letter Queues and PostgreSQL state logs eliminate the need to manually monitor crashed browser sessions.
- Future-Proof Extensibility: Decoupling orchestration from browser logic allows the client to plug in new municipal portals and compliance workflows without rebuilding core infrastructure.
Conclusion
Transforming manual certificate retrieval into an automated AWS pipeline addressed a critical operational challenge while establishing a scalable framework for future growth. By replacing manual portal navigation with an asynchronous, event-driven platform, the organization eliminated a major administrative bottleneck and laid the groundwork for long-term operational efficiency. This architecture delivers a resilient, cost-effective foundation that enables the seamless integration of new regional compliance workflows as business demands evolve.

