TGViewer
Dateno Dateno @datenosearch · 105 subscribers
Post #9 101
🔍 One of the key features of the Dateno search engine is that, in addition to collecting basic metadata about datasets and APIs, its crawlers also gather links to related resources and even archive some of them.

This approach not only helps provide users with a convenient tool for finding data but also allows us to analyze how data is actually published and in what formats.

📊 As of July 2025, Dateno has indexed 5,961,849 datasets from open data portals. That’s about 27% of the total datasets, map layers, and time series aggregated from data catalogs, geoportals, and statistical databases.

Let’s take a closer look at these 5.9 million datasets:
Some datasets come without any associated files, while others may include dozens or even hundreds of attached resources. That’s why, when analyzing file formats, it makes more sense to focus on the number of resources rather than the number of datasets.

👉 Currently, Dateno indexes 6.7 million resources (files and links) attached to these datasets—on average, around 1.1 resources per dataset.

Here’s the breakdown of the most common file formats:

CSV: 1,008,646 files (15%)

XLSX: 525,329 files (7.8%)

XML: 522,501 files (7.8%)

JSON: 509,668 files (7.6%)

ZIP: 496,709 files (7.4%)

PDF: 487,189 files (7.3%)

HTML: 475,377 files (7.1%)

WMS (geospatial API): 320,159 files (4.8%)

NC (NetCDF): 233,229 files (3.5%)

XLS: 185,855 files (2.8%)

WCS (geospatial API): 141,472 files (2.1%)

KML: 122,781 files (1.8%)

DOCX: 115,723 files (1.7%)
…and many more.

It’s no surprise that CSV remains the most popular format for open data publication. Other common formats include XLSX, XML, JSON, and legacy XLS files.

Formats like WCS, WMS, and KML reflect the increasing role of geospatial data published via standardized APIs and file formats.

Meanwhile, the popularity of PDF, DOCX, and HTML points to the reality that not all datasets come with machine-readable files. Sometimes data is shared as reports, documents, or links to external sources, requiring additional effort to extract the actual data.

📉 And what about data science-friendly formats?
Take Parquet files, for example—widely used in data engineering and data science for their efficiency. Surprisingly, only 1,652 Parquet files are currently indexed by Dateno—less than 0.025% of all resources. Quite an eye-opener!

The world of open data is still far from being fully aligned with the needs of data engineering and data science. Closing this gap is essential if we want to unlock the full potential of open data for advanced analytics and AI.

#OpenData #DataScience #DataEngineering #Dateno #DataFormats #Parquet #CSV #GeospatialData #AI
  • ✍ 3
  • 🔥 3
  • ❤ 1
More from @datenosearch
  1. Feb 3, 2026New at Dateno: Python SDK, MCP Server, and What’s Coming Next We started the year with sev…
  2. Dec 19, 2025We’ve launched Dateno API v2 -- a major upgrade to our data search platform We’re excited…
  3. Dec 8, 2025Open Data in Armenia: No National Data Portal - Yet One of the most notable characteristic…
  4. Dec 2, 2025Publishing an example of AI-enhanced open data: Internacia Datasets — open-source and open…
  5. Nov 22, 2025🚀 Major Update of the Dateno Data Catalog Registry The Dateno Registry — an open-source &…
  6. Nov 17, 2025Regular country open data overview, this time Estonia — Open Data in Estonia: A Small Coun…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →