RatingAdda
RatingAdda
Ratings with the receipts
Home / Methodology

Methodology

How RatingAdda gathers, filters, scores, and ranks products.

Last updated: June 1, 2026

RatingAdda ranks products from real public Reddit conversations, not from press releases or paid placements. Every step below is automated; the only human in the loop is admin curation for obvious mistakes (a comment about something else getting attached to the wrong model, for example).

1. Where the data comes from

We use the public Reddit API to read posts and their comment threads in the subreddits where buying advice for a category is actually discussed (e.g. r/headphones, r/HeadphoneAdvice for headphones). We do not access private subreddits, deleted content, or anything that requires a Reddit login.

2. Filtering for relevance

Not every Reddit post in a category subreddit is a genuine recommendation discussion. We drop posts that are off-topic, meta-analyses (someone summarising other people's ranks), joke threads, or low-quality replies (single emoji, “same”, etc.).

3. Extracting the product mention

An LLM reads each comment and extracts the specific product being discussed, resolving shorthand like “XM5” or “HD600” to the full canonical name (Sony WH-1000XM5, Sennheiser HD 600). Comments that don't name a specific model don't count.

4. Sentiment toward the product

We don't judge a whole comment with one sentiment label, because a comment can praise one product while criticising another. Instead we measure the sentiment aimed at the specific product named in the comment (aspect-based sentiment). Each judgement carries a confidence score, and only high-confidence opinions count as positive or negative; anything uncertain is treated as neutral. This is what stops a comment praising a rival product from being filed under this one. You can see the result on every category page: clauses that are positive about the product are underlined in green, negative ones in red.

5. Per-user dedup, then ranking

If one Reddit user wrote five comments about the same product, that's one vote, not five. Their stance is the majority of their clauses for that product (ties go to the most-upvoted one). The headline number on every product card is “X% positive of N users”. Here N is the number of unique users who expressed a clear positive or negative opinion about the product (not everyone who merely mentioned it), and X% is the share of those users who were positive. Total discussion volume is shown separately as the mentions count.

Our default ranking ("Best Overall") combines two signals:

  • 75% volume of positive opinion: how many unique users said something positive, relative to the most-loved product in the category. Rewards products people actually keep recommending, not just niche favourites.
  • 25% positive:negative ratio: how clean the praise-to-criticism ratio is, ignoring neutrals. Stops a product from coasting on volume alone if it also has a large complaint pile.

Two other tabs let you re-sort: Most Discussed (raw mention count) and Highest Approval (highest %, with a minimum of 5 users so a one-comment fluke doesn't top the chart).

6. India-availability and specs

We surface only products that are sold in India. The check is silent - we don't put an “available in India” badge anywhere because every product on the site already passed it. Specs (driver size, ANC, weight, etc.) are pulled in once per product and editable by admins if the auto-extraction got something wrong.

7. What we don't do

  • We don't average star ratings from other sites.
  • We don't display prices or track price drops.
  • We don't accept payment to feature, rank, or recommend a product.
  • We don't claim to find every product in a category - only those with enough Reddit discussion to draw a confident conclusion.

Limitations

  • Reddit bias. Reddit skews toward enthusiasts. A product can be unfairly maligned (or hyped) by a vocal subset.
  • Sarcasm. Sentiment models miss obvious sarcasm. We try to filter the worst offenders but some leak through.
  • Time. Old products keep accumulating mentions. We weight recent discussion slightly higher but a 2017 classic with 2,000 mentions can still beat a 2024 release with 200.
  • Categories. We launch with a small number of categories so the data is dense per product. Niche models get cold-started over time.

Data freshness

The pipeline runs weekly. Each category page shows the last update timestamp at the top. If a product just launched and isn't on the site, give it a few weeks of Reddit discussion to accumulate, or use the request-a-category form to ask us to expand coverage.

Questions or corrections? Email hello@ratingadda.com.