Data methodology
From catalogue page to comparable record.
CredScape turns public continuing education information into structured records you can search, compare, and trace. Every record retains its source and collection timestamp.
Credential language resolves to CTDL.
Institutions use different names for similar credentials. Every offering maps to a published Credential Transparency Description Language type, which makes one provider's certificate comparable to another's.
Three sources, one boundary.
Direct institutional crawls carry the records that power market composition. Platform catalogues broaden discovery and stay out of composition statistics.
Delivery mode, normalised.
Online, hybrid and in-person resolve to the same three values across every provider, so a delivery filter means the same thing in every market.
Prices you can put side by side.
Prices normalise to a common basis, so a median or a band means the same thing whether the offering is priced per participant, per seat, or per cohort.
Collection process
A clear path from source to record.
Four stages make every transformation visible.
Public catalogues enter in batches, fields are checked, source language resolves to a shared vocabulary, and the finished record receives its collection date.
Collect
Capture the offering page, source URL, and published details.
Validate
Check the fields available from that source before publication.
Resolve
Map delivery, credential, price, and duration into common fields.
Date
Attach the batch collection timestamp to the finished record.
Normalisation
Source detail stays intact. Meaning becomes comparable.
Institutions describe similar offerings in different ways. CredScape preserves the published detail, then resolves the fields that drive comparison.
Delivery becomes online, hybrid, or in person. Price and duration retain their component parts, so different presentation styles can be read on the same basis.
Validation rules
Rules that protect meaning.
Quality is enforced field by field.
Each rule keeps a source fact interpretable after it joins records from other providers.
Keep the source attached
The source URL, source type, and collection timestamp stay with the offering record.
Applied on writeLeave gaps visible
An unavailable field stays unreported, preserving the boundary of the published source.
Applied on writePreserve price structure
Amount, currency, cost type, payment pattern, and free status remain separate fields.
Applied on writePreserve duration structure
Exact, minimum, maximum, and hour-based duration details remain available when published.
Applied on writeResolve credential language
Institutional terms map to the Credential Transparency Description Language vocabulary.
Applied on writeDate the collection
Batch collection gives each record a timestamp for record-level recency.
Applied on writeRetrieval
Search the way the question is asked.
Every offering receives a dense vector embedding. Semantic similarity and full-text search work together, then filters narrow the result by geography, platform, credential form, and delivery mode.
Credential vocabulary
One classification system across every provider.
CTDL creates a common language for credentials.
Source terminology remains available while the resolved form supports consistent filtering and analysis.
Corpus scope
Two source roles. One transparent boundary.
Direct institutional crawls across Canada and the United States power market-composition dashboards. The Coursera and edX catalogues expand searchable discovery.
Platform subscription pricing sits outside market-composition statistics so price medians retain a consistent basis.
Reading the record
Missing fields remain meaningful.
An empty skill, duration, or price field means the structured source lacked that detail. The collection timestamp marks the batch date. Current page contents may differ after collection.
Put the methodology to work on your market.
Bring a program question. We will show you the records, sources, and comparisons behind the answer.