An illustrative view of how repository, organization, rel…
.
Data Marketplace · Markets & technology
Map public GitHub repositories, organizations, releases, languages, topics, and activity for software ecosystem research.
.
{
"repository": "example-org/example-project",
"repository_url": "https://example.test/repository",
"description": "Illustrative repository record",
"languages": [
"JavaScript",
"Python"
],
"topics": [
"data"
],
"latest_release": {
"tag": "v2.4.0"…Schema and fields
Start with repository, organization, release. Source identity, collection time, and field meaning stay connected so records can move into analysis without losing their context.
repository stringOwner and repository identity
repository_url urlPublic source URL
description stringPublic repository summary
languages arrayObserved language context
topics arrayPublic repository topics
latest_release objectAvailable public release context
activity objectObserved public activity signals
captured_at timestampCollection time
Records in this collection
Explore the GitHub on-demand API
Repository
Organization
Release
Language
Activity observation
Freshness and delivery
Use a dated ecosystem snapshot or recurring refreshes when repositories, releases, languages, and public activity need to be compared over time.
Snapshot Start from a dated baseline Selected scope and collection window
Refresh Receive recurring updates Cadence and change behavior
Format Use JSON, CSV, or Parquet Schema, partitioning, and manifest
Validation Run repeatable checks Counts, fields, missing states, and version
GitHub applications
Use source-linked, time-stamped records for technology ecosystem mapping, plus open-source project research. Each application starts from the same documented fields and collection context.
01
Compare public repositories, organizations, topics, languages, and release patterns.
02
Build source-linked project records for selected technologies, organizations, or categories.
03
Track newly observed releases and public repository activity across a controlled scope.
From sample to production
Start with representative rows, run the joins and calculations that matter, and shape delivery around the system that will consume the data.
01
Select the entities, markets, fields, dates, and business output the collection should support.
02
Check identifiers, fields, joins, missing states, and source context with your own queries or models.
03
Choose the snapshot or recurring schedule, file format, partitions, manifest, and destination.
GitHub dataset FAQ
Map public GitHub repositories, organizations, releases, languages, topics, and activity for software ecosystem research. Core records include repository, organization, release, language, activity observation.
Common applications include technology ecosystem mapping, open-source project research, release and activity monitoring. Start with the fields and time window required by the business output.
Use a dated ecosystem snapshot or recurring refreshes when repositories, releases, languages, and public activity need to be compared over time.
Use the on-demand API when your application chooses individual requests and timing. Choose this dataset for a prepared bulk snapshot or recurring file delivery.
Continue the data journey
Compare the adjacent dataset, open the source API, or continue into the relevant data, industry, and solution pages.
Start with real records
Tell us which repository and organization data you need, plus the markets, fields, dates, and destination. We will map the collection to that requirement.