The challenge
Consumer wearables record activity and sleep every day, which makes them an attractive source for studies. Turning that stream into research data is harder than it looks. Every provider has its own sign-in flow, payload format and rate limits. OAuth tokens expire, and when one does, data simply stops arriving with no visible error, leaving a gap that may only come to light much later during analysis.
Timing is the second problem. Providers send webhooks on their own schedule, and they often resend or revise a day’s figures. If the web application processes all of that inline, a busy period slows the very screens participants use to connect their devices, which is the worst possible moment to lose them. REDCap, meanwhile, expects one tidy record per subject per day, so a stream of partial and repeated updates has to be reconciled before it reaches the study database.
- Let participants connect devices without a manual, error-prone setup
- Stop expired tokens from silently interrupting data collection
- Absorb webhook bursts and rate limits without slowing the participant app
- Deliver one consolidated daily record per subject to REDCap Cloud, with a reviewable trail
Our approach
We began by separating two very different workloads. The participant-facing Flask application needs to be quick and simple, so it handles onboarding and little else. Device connections go through the Terra API, which links participants’ Fitbit, Strava and other provider accounts through secure widget sessions. Using one aggregator saves the client from building and maintaining a separate integration for every wearable brand.
We treated token expiry as a certainty. OAuth refresh runs automatically and restores connections on its own, because in a study a missing token means missing data, and nobody should have to notice the gap by hand.
Incoming webhooks are normalised into PostgreSQL, so readings from different devices share one consistent structure. All heavy sync and stream processing is handed to Celery workers backed by Redis, away from the Flask application thread. That lets the work fan out asynchronously, lets us pace calls to stay inside each provider’s rate limits, and keeps the onboarding screens responsive however much data is arriving.
On the research side, scheduled jobs deliver consolidated daily metrics into REDCap Cloud (RCC), at an interval the team configures. We made every write safe to repeat. Merges are UPSERT-friendly, so each subject’s activity and sleep data for a given day always resolves to one record, and a late correction updates that record instead of adding another. Monitors can download CSV exports for review, and audit logs capture webhook and API activity, giving a trail of what came in and what went out. Participants, for their part, connect devices through a responsive glass-style screen.
- Terra API onboarding with automatic OAuth token refresh
- Webhook data normalised into PostgreSQL
- Celery and Redis workers kept apart from the Flask app
- Daily merges into REDCap Cloud on a schedule, safe to repeat (UPSERT)
- Audit logs of webhook and API traffic, plus CSV exports for monitors
- Processing paced for provider rate limits, with async fan-out
Architecture & stack
- Application
- Python, Flask
- Data
- PostgreSQL, Redis
- Background processing
- Celery, Schedulers
- Integrations
- Terra API, Fitbit, Strava, Webhooks, OAuth, REDCap Cloud (RCC)
- Reporting & audit
- CSV exports, Audit logs
The outcome
The study team receives wearable data in the system it already works in: one consolidated record per subject per day in REDCap Cloud, CSV exports whenever monitors need them, and an audit trail of what arrived and what was passed on.
Because the heavy processing runs in background workers, busy periods and provider limits are absorbed without participants noticing. Automatic token refresh keeps devices connected, and repeat-safe merges stop resent or revised data from creating duplicate records in REDCap.
For any client connecting health devices to a research workflow, our advice is the same. Assume every third-party connection will fail at some point and design the recovery in. Process incoming data in the background, make every write idempotent, and log each exchange, because sooner or later someone will ask where a number came from and you will want the answer on record.
Last updated



