<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Visualization on Eigenform Articles</title><link>https://www.eigenform.ai/insights/tags/data-visualization/</link><description>Recent content in Data Visualization on Eigenform Articles</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Mon, 10 Aug 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://www.eigenform.ai/insights/tags/data-visualization/index.xml" rel="self" type="application/rss+xml"/><item><title>How Reports From The Future Draws Its Star Maps</title><link>https://www.eigenform.ai/insights/how-reports-from-the-future-draws-its-star-maps/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0800</pubDate><guid>https://www.eigenform.ai/insights/how-reports-from-the-future-draws-its-star-maps/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Reports From The Future maps AI-related vocabulary as it emerges, aiming to spot &amp;ldquo;semantic basins of attraction&amp;rdquo; - ideas circulating in tech discourse before they have a settled name - across six months of arXiv papers and three months of web discourse (X, HuggingFace blogs, Hacker News, LessWrong).&lt;/li&gt;
&lt;li&gt;All three plates share a pipeline: documents are embedded into vectors, UMAP reduces them to a drawable space (fit once and frozen, so the sky doesn&amp;rsquo;t scramble week to week), and HDBSCAN clusters dense regions into constellations, leaving outliers as unassigned &amp;ldquo;dust&amp;rdquo; rather than forcing them into a group.&lt;/li&gt;
&lt;li&gt;Plate I (the Research Sky) names constellations with TF-IDF only, and corrects each month&amp;rsquo;s colour for publishing rate rather than raw count - a fix that moves August from 6% to about 70% of the May peak once partial and weekend-heavy days are properly weighted.&lt;/li&gt;
&lt;li&gt;Plate II (the Attention Sky) runs HDBSCAN twice, once over a rolling 91-day background for context and once per week at a smaller threshold for the constellations actually drawn, and names clusters with raw TF-IDF terms on the view that a badly-formed name is itself information about unsettled vocabulary.&lt;/li&gt;
&lt;li&gt;Plate III (the Emergent Sky) reuses Plate II&amp;rsquo;s clusters but adds a &amp;ldquo;name gap&amp;rdquo; metric - the percentile-rank difference between semantic and lexical coherence - to separate genuinely emerging concepts (high shared meaning, low shared vocabulary) from buzzwords (high shared vocabulary, low shared meaning), each routed through a different, deliberately inverted naming prompt.&lt;/li&gt;
&lt;li&gt;The naming pipeline for Plate III runs four models independently to describe each cluster, then has one model synthesise a name from where their descriptions converge, and a final model distil that into a short phrase - illustrated with a real example of &amp;ldquo;Claude&amp;rdquo; being described mid-genericide as a term coming apart rather than settling on a single referent.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Takeaways&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>