🧠 Blindspot β€” Find What You Don't Know You're Missing


Every day you do this:

You open Twitter. You skim some papers. You read a blog post. You bookmark a few things. You feel caught up. But are you?

The problem is: you only find what you already know to search for.

The concepts that would actually change your research direction? The ones just outside your radar? The ones your peers adopted 6 months before you?

Those are your blindspots. And you'll never find them by searching.


πŸ’‘ What Blindspot does

Blindspot is an AI trained with SFT (Supervised Fine-Tuning) to act like a research assistant that:

Step What happens
Reads your profile Understands your current work and past papers
Scans 1,168 ML concepts All in 3ms β€” what you'd take weeks to browse manually
Decides what to surface Not just trending β€” specifically what you will adopt and understand
Shows you the proof Every recommendation checked against real adoption ground truth

πŸ”¬ The RL training explained simply

The AI learned by doing what you do β€” inspect a topic β†’ decide to keep or skip β†’ stop when done.

Reward signal:

  • βœ… +1 if you would have adopted it
  • πŸ“ˆ +0.5 if it improves your understanding
  • πŸ”₯ +0.5 if it's novel to you specifically
  • ❌ βˆ’0.1 for every useless recommendation

After 3,200 training episodes across 13 real ML researchers: the AI learned to beat every baseline β€” including the "just show trending" approach.


πŸ‘‡ Start with Tab 1: pick a researcher and compare the results

Step 1 β€” Pick a researcher

These are 17 actual researchers in our database (we tracked what concepts they adopted over time).

What you'll see:

  • 5 different strategies tried on the same person
  • Each strategy scored: did they actually adopt what was recommended?
  • The key comparison: AI before training vs AI after SFT training

πŸ’‘ Start with the default user, click Run, and scroll down

πŸ‘€ Pick a researcher

The full picture


πŸ”„ What happens when you click "Find my blindspots"

Your profile / paragraph
        ↓
Matched to closest researcher in our database (TF-IDF cosine similarity)
        ↓
5 strategies run on their 40-concept candidate pool:

    Random        β†’ picks 3 at random
    Trending      β†’ picks 3 most popular
    Dense         β†’ picks 3 most similar to your past work
    Pre-training  β†’ Qwen2.5-1.5B base model picks (no SFT)
    SFT trained   β†’ Qwen2.5-1.5B + LoRA (SFT on 40 expert traces) ← this is Blindspot
                ↓
Each strategy scored:
  βœ… Did the researcher actually adopt this concept? (+reward)
  πŸ“ˆ Did it improve their understanding? (+reward)
  πŸ’Ž Was it a non-obvious pick? (+reward)
  ❌ Was it a waste of time? (βˆ’reward)
        ↓
Results shown side-by-side with before/after toggle

πŸ“Š Real calibration numbers (5 seeds Γ— 17 researchers)

Strategy Mean reward What it means
Random βˆ’0.01 Noise β€” proves reward is calibrated
Trending +1.11 Good, but not personalized
Dense Retrieval +0.41 Relevant, but obvious picks
Blindspot (before RL) βˆ’0.47 Base model struggles
Blindspot (after SFT) +1.85 RL learned what each person needs
Oracle (upper bound) +2.77 What perfect knowledge would score

πŸ—οΈ Architecture

  • Training: Qwen3.5-9B + LoRA via Unsloth, trained with TRL's SFTTrainer (3 epochs Γ— 40 traces, H100)
  • This demo: Zero GPU β€” all trained-policy responses pre-cached in data/demo_cache.json
  • Data: 17 real ML researchers, 1,168 concepts, 282 reading paths, 62 adoption pairs
  • Held-out test: 4 researchers never seen during training

Code: github.com/vasarlalikhilavinash/blindspot-env

Trained adapter: huggingface.co/Vasarlaavinash/blindspot-sft-1.5b