Builds

DevAtlas Maps 30,000 Dev Tools So You Stop Digging Through Awesome Lists

· Builds

Awesome lists rot. I built DevAtlas to treat them as data sources and re-index 30,000+ tools daily — here is how the search layer works.

DevAtlas Maps 30,000 Dev Tools So You Stop Digging Through Awesome Lists

Last updated: October 3, 2026 · 4-minute read

The developer universe has a map. It is 4,000 markdown files called "awesome-something", and it is mostly wrong. Lists get famous, get starred, then get abandoned; entries point to dead repos; the tool you actually need is sitting in a niche list nobody linked from the front page. DevAtlas is my answer: treat every awesome list as a data source, normalize it into one index, and refresh it daily. It has been live at devtomb.netlify.app with over 30,000 tools, APIs and datasets, and this post is the behind-the-scenes.

The data model beats the UI

The naive version of DevAtlas is a search box over concatenated markdown. That version dies the first week, because raw awesome-list entries are inconsistent soup: one list writes Tool - description, another uses tables, a third hides descriptions in headers. The entire build hinges on a normalization pass that collapses all of it into one record shape — name, description, category, source list, repo status.

The index then answers the two questions awesome lists cannot. First, what exists — search across every list at once, not one curated silo. Second, what is alive — because each record keeps its source, stale sources can be weighted down and dead links can be pruned instead of discovered one frustrating click at a time.

Why daily refresh matters more than it sounds

A tools directory is a perishable good. Look at any "best tools 2024" list and count the dead repos — decay is measured in months. A daily re-index turns DevAtlas from a snapshot into a live map: new tools surface without a maintainer deciding they are worthy, and abandoned ones fade instead of polluting results. The refresh runs unattended; if a source list dies, the index keeps working off the rest.

There is an honest trade-off: automated aggregation means occasional duplicate entries and imperfect categorization. I accept a small number of imperfect rows over a large number of stale ones. Curation ages like milk; indexing ages like wine.

Search that respects the long tail

The UI is deliberately boring — a fast search box, category facets, result rows. The interesting work is in ranking. A tool with 40,000 stars and a tool with 200 stars can both be exactly right depending on the query, so popularity is a tiebreaker, not the sort key. Text match quality and category fit come first. That single decision is why the long tail is actually findable: search for a niche need and you get the niche tool, not ten JavaScript frameworks that mentioned the word once.

For discovery beyond search, browsing by category works the way you would hope — and when the need is specifically "an API for my prototype", that is a job for APIVault, which narrows the same philosophy to APIs with auth and CORS filters. The two tools started as one project and split, and they are better siblings than they ever were as a monolith.

Inside the pipeline

The system has three stages, and each one taught a lesson. Ingestion pulls the awesome-list corpus and applies the normalization pass — a stack of parsers that each handle one list format, with a fallback that extracts link-plus-description pairs when no structured parser matches. Resolution deduplicates: the same tool surfaced by five lists becomes one record with five source attributions, which is how the index stays honest about popularity without double-counting. Serving is the boring part on purpose — a prebuilt index, static-friendly queries, no database server between the user and the answer.

The failure archive is instructive. Lists that rename themselves mid-cycle, entries whose description lives in the commit message, tools that changed names and kept their old listing alive — every one of these broke the naive version first and got a rule later. An index of 30,000 records is really a list of 30,000 exceptions wearing a uniform, and the daily refresh is what keeps the exceptions from accumulating into garbage.

What I learned building it

Three lessons transfer to any aggregation project. Normalize early, decorate late: every feature that ever worked in DevAtlas works because the record shape is clean, and every bug that ever annoyed me lived in the ingestion edge cases. Freshness is a feature users can feel: nobody writes a review saying "the data was from today", but stale directories die from exactly that silence. Boring stacks ship: React, TypeScript and a static-friendly index — nothing in DevAtlas needs a database cluster, which is why it can live on Netlify and stay free.

The next increment on my list is repo-health signals beyond link-checking — commit recency and license surfacing, so the map also tells you which tools are safe to build a business on. If that sounds useful, the live site is devtomb.netlify.app, and the source of my other experiments is the lab. For the philosophy behind tools like this, the open-source first-principles guide covers why I build free discovery tools at all.

---

Not affiliated with GitHub or the awesome-list maintainers aggregated here. Sources: the DevAtlas site, its own documentation, and the awesome-list ecosystem it indexes.

ansaribilal.com — technology, tested in public.