feat: integrate Google Scholar alerts and improve paper filtering pipeline

- Add Google Scholar Alert collection via Gmail IMAP
- Add Gmail App Password based authentication
- Add Google Scholar environment variables to NAS compose configuration
- Enable Google Scholar Alert source in config.nas.yaml
- Disable subject filtering that incorrectly excluded Scholar alert emails
- Parse paper titles and links from Google Scholar alert HTML emails
- Deduplicate collected Scholar papers by normalized title
- Filter Scholar UI/control links such as update alert and unsubscribe entries
- Filter bracket-only alert labels such as [automotive radar]

- Add Google Scholar specific relevance threshold
- Keep global min_relevance at 2
- Set Google Scholar minimum relevance to 3
- Reduce Scholar candidates from 229 collected / 207 merged to 28 accepted

- Add Gemini API rate limiting
- Enforce minimum 13 second interval between Gemini requests
- Apply shared rate limiter to abstract and full-text enrichment
- Prevent Gemini free-tier 5 RPM quota errors
- Verify retry processing of previously failed AI enrichment jobs

- Verify end-to-end NAS workflow
- Google Scholar Alert collection successful
- Gemini enrichment successful without HTTP 429 errors
- Joplin report generation and WebDAV synchronization successful
This commit is contained in:
2026-08-21 00:25:24 +09:00
parent c3aad03182
commit 2543b0c23f
8 changed files with 233 additions and 51 deletions
+14 -1
View File
@@ -259,7 +259,20 @@ def run(config_path, dry_run=False):
merged[k] = merge_paper(merged[k], p) if k in merged else p
papers = [classify_and_score(p, search_cfg) for p in merged.values()]
papers = [p for p in papers if p.relevance >= int(app.get('min_relevance', 1))]
min_rel = int(app.get("min_relevance", 1))
scholar_min_rel = int(
app.get("google_scholar_min_relevance", min_rel)
)
papers = [
p for p in papers
if p.relevance >= (
scholar_min_rel
if p.source == "Google Scholar Alert"
else min_rel
)
]
db_path = resolve_path(cfg, app.get('database_path', './data/papers.db'))
db = PaperDB(db_path)