Data collection ยท weienwong.online
Reddit datasets, on a schedule
Describe what you want collected โ a keyword, a subreddit, a username, a date window โ and this service keeps pulling it for you. Results are deduplicated, stored, and exportable, so you get a dataset rather than a pile of screenshots.
One account works across all thirteen Weien Wong services.
What it does
Filtered collection
Keyword, subreddit, username and date-window filters, combined however you need them.
Recurring schedules
A job re-runs on the interval you set, so the dataset grows without you touching it.
Deduplicated records
Posts already collected are not collected twice, however many times a job runs.
Export when you need it
Browse results in the app, or pull the whole set out for analysis elsewhere.
Getting started
about two minutes- Create an account. Use the button above โ it opens right here, no redirect. The same account signs you in to every other service on the domain.
- Create a scrape job. Give it a keyword or subreddit and a time window. Start narrow; you can widen it once you see what comes back.
- Set the schedule. Choose how often it re-runs. Records accumulate across runs, deduplicated.
- Read or export. Open a record for the full post data, or export the set once the job has collected enough.
Before you start
This collects publicly available posts. You are responsible for how you use what you collect, including Reddit's terms and any data-protection rules that apply where you are. Collection is rate-limited on purpose โ a job that looks slow is usually a job being a good citizen.