openskills.info
Course Preview

Blockchain Indexing and Subgraphs

Blockchain indexing turns raw chain activity into queryable data. Subgraphs define which contracts, events, and entities to index, and The Graph Network turns those definitions into GraphQL APIs applications can query without running their own nodes.

itDistributed systems, messaging, and integration

Recommended first:web3-application-integration

Don't Panic: Blockchain Indexing and Subgraphs

A blockchain remembers transactions. It is considerably less helpful when an application asks for a tidy list of everything a particular contract did last week. The chain exposes blocks, logs, and state through RPC. It does not hand you application records with filters, joins, and convenient names. You can collect the raw data yourself, but then you have built an indexer, whether or not you meant to.

A subgraph is a recipe for that indexer. Its manifest names the chain, contracts, and events to watch. Its schema names the records an application wants to query. Its mappings turn each matching event into one or more stored records. Graph Node follows the recipe and serves the result through GraphQL. That is the whole path: watch, transform, store, query. The files are small; the historical chain they may have to process is not.

The first surprise is that the schema is an operational choice. If a token has many transfers, copying every transfer ID into one growing token array seems natural. Each new event then rewrites a larger array. Store the token reference on each transfer instead, and let @derivedFrom expose the collection from the token. The relationship is still queryable without asking one entity to carry the entire history.

The second surprise is that a successful GraphQL response can be old news. A subgraph must first catch up with the chain. Its _meta field reports the latest block it has indexed and whether indexing errors exist. Compare that block with the chain head when freshness matters. A dashboard that loads is not necessarily a dashboard that has caught up.

Subgraph Studio is the workbench. A staged deployment is for building, testing, and checking logs. Publishing is a separate onchain step that makes the subgraph available on The Graph Network, where Indexers serve queries. An application can still use direct RPC for a current contract value or to submit a transaction. The subgraph supplies the historical and cross-contract view that raw RPC does not arrange for you.

If the ordinary mapping path is too slow for a large backfill or a streaming transformation, Substreams offers parallel processing over block streams. It can send results to a database or feed a subgraph when GraphQL is still the right public interface. The distinction is practical: decide what data the application must ask for, then choose the processing and query path that can deliver it.

Read the Intro for the full data path and network roles. Use the Cheatsheet when choosing handler types, entity relationships, and query pagination. The Practice Reference and Exercise turn those decisions into a staged subgraph you can inspect before anyone depends on it.

Where this skill leads

Relevant careers

See how this topic contributes to broader role-level skill maps.

Sources