<!-- llms-explorer concept facts · https://llms-explorer.com/tree/data-acquisition-and-sampling/ · pack 2026-09-08 · ~6161 tokens -->

# Data Acquisition and Sampling

> This skill covers the third stage of the data analysis curriculum: getting data into a form the

Parent: [Data Analysis](https://llms-explorer.com/tree/data-analysis/) · 33 facets · 83 facts · page: https://llms-explorer.com/tree/data-acquisition-and-sampling/

## Data Acquisition and Sampling

- This skill covers the third stage of the data analysis curriculum: getting data into a form the analysis can operate on, and constructing a sample that supports valid inference about the target population. The two activities are intertwined: the choice of source constrains what sampling design is possible, and the sampling plan determines which sources are acceptable. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#data-acquisition-and-sampling)
- A common mistake is to treat acquisition as a logistics problem - "just pull the data" - and discover only at the analysis stage that the population was wrong, the sample frame had coverage gaps, the schema drifted mid-pull, or the file format made the planned query infeasible. This stage owns the responsibility for catching these failures before they contaminate downstream work. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#data-acquisition-and-sampling)

## Sub-skill routing table

- This hub absorbs 9 former standalone skills as on-demand reference files. When a task matches a row, Read the listed references/ file before answering - do not rely on this table alone for depth. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#sub-skill-routing-table)

## 1. Data Source Taxonomy

- Three orthogonal axes characterize any source. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#1-data-source-taxonomy)

## 1.1 Primary vs secondary

- Primary data is collected directly to answer the current question (surveys you designed, interviews, sensor data, A/B exposure logs). Secondary data was collected by someone else for a different purpose and is reused (Census/BLS, commercial panels, academic archives, third-party API exports). Strong analyses combine both. The trade-off is purpose-fit (primary) versus speed/scale/cost (secondary). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#11-primary-vs-secondary)

## 1.2 Structured vs semi-structured vs unstructured

  - Structured: row/column tabular with fixed schema (relational DBs, warehouses, CSV/Parquet). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#12-structured-vs-semi-structured-vs-unstructured)
  - Semi-structured: hierarchical/self-describing (JSON, XML, logs, NoSQL docs, Avro/Protobuf). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#12-structured-vs-semi-structured-vs-unstructured)
  - Unstructured: free-form text, images, audio, video; requires feature extraction (OCR, ASR, embeddings) or a model interface. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#12-structured-vs-semi-structured-vs-unstructured)
- Vector embeddings and LLMs reduced the cost of operating on unstructured data, but the closer a source is to structured form, the cheaper and more deterministic the analysis. Schema-on-read (data lakes) defers structure to query time; schema-on-write (warehouses) enforces it at load time. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#12-structured-vs-semi-structured-vs-unstructured)

## 1.3 Internal vs external

- Internal sources (production DBs, event logs, CRM, telemetry) are more reliable and granular but may not generalize. External sources (APIs, public datasets, scraped pages, panels) broaden the population but add coverage uncertainty, licensing risk, and schema drift. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#13-internal-vs-external)

## 2.1 REST, GraphQL, gRPC

  - REST (HTTP+JSON): broadest compatibility, HTTP caching, OpenAPI contracts. Downside: over/under-fetching. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#21-rest-graphql-grpc)
  - GraphQL: client asks for exactly the fields it needs; eliminates over/under-fetch. Downside: HTTP caching is harder, rate limiting via query-cost budgets, per-field authorization. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#21-rest-graphql-grpc)
  - gRPC (HTTP/2 + Protobuf): code-generated clients, multiplexed streaming, 3-10x smaller payloads. Best for internal service-to-service traffic. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#21-rest-graphql-grpc)
- Common 2026 pattern: REST public, GraphQL BFF/frontend, gRPC internal. For acquisition you mostly meet REST and GraphQL. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#21-rest-graphql-grpc)

## 2.2 Authentication

  - API key: shared secret in a header; identifies the app, not a user; TLS only; rotate. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#22-authentication)
  - OAuth 2.0: access + refresh tokens. Authorization Code with PKCE for user apps; Client Credentials for service-to-service. Validate scopes server-side. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#22-authentication)
  - JWT: signed bearer token; stateless verification; keep short-lived; verify the alg header (avoid alg: none). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#22-authentication)
  - mTLS / certificate auth: high-trust internal/financial/healthcare APIs. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#22-authentication)
- Treat refresh tokens as the most sensitive secret: encrypt at rest, rotate on suspicion, log every refresh. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#22-authentication)

## 2.3 Pagination

  - Offset/limit and page number: simple but break under concurrent writes. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#23-pagination)
  - Cursor-based (opaque token): stable under writes; preferred for high-volume APIs. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#23-pagination)
  - Keyset/seek (sort key + tiebreaker): cheap on indexed columns. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#23-pagination)
- Persist the cursor after every page so a partial failure can resume. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#23-pagination)

## 2.4 Rate limiting

- Fixed window, sliding window, token bucket, cost-based (GraphQL). Build clients with exponential backoff on 429/503, respecting Retry-After. Cap retries. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#24-rate-limiting)

## 2.5 Webhooks

- Inverse of polling. Always verify the signature header (HMAC-SHA256 over the raw body) before trusting the payload; treat unsigned webhooks as untrusted input. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#25-webhooks)

## 3. Web Scraping

- Acquisition without a contract. Use only when there is no API and the legal/ethical posture is sound. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#3-web-scraping)

## 3.1 Tooling

- requests + BeautifulSoup: static HTML, no JS. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#31-tooling)
- Scrapy: full crawling framework (concurrency, throttling, retries, pipelines). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#31-tooling)
- Playwright / Selenium: headless browsers for JS-rendered or auth-gated pages; slower. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#31-tooling)
- TLS-impersonation tooling (curl_cffi): a signal the site does not want to be scraped. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#31-tooling)

## 3.2 Legality

- As of 2026 the hiQ v. LinkedIn line: scraping publicly accessible data generally does not violate the CFAA. But the full risk surface includes CFAA (bypassing auth/access controls), Terms of Service (civil claims), copyright (bulk reproduction), GDPR/CCPA (personal data, stricter in the EU), and EU database rights. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#32-legality)

## 3.3 Ethics

- Respect robots.txt; set a descriptive User-Agent with contact info; throttle (≤1 req/sec for small sites); cache aggressively; avoid PII without a lawful basis; never bypass authentication, paywalls, or rate-limiting controls. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#33-ethics)

## 4.1 Bulk export

- SELECT * into CSV/Parquet/Avro. Simple for cold historical data; anti-pattern for large operational tables. Use native utilities (mongoexport, pg_dump, mysqldump --single-transaction, bq extract, Redshift UNLOAD). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#41-bulk-export)

## 4.2 Incremental polling (JDBC/ODBC)

- Connectors (Kafka Connect JDBC, Airbyte, Fivetran, Meltano) poll on a schedule, identifying changes by updated_at or auto-increment id. Gaps: hard deletes invisible, backdated updates missed, high-frequency polling stresses the source, schema changes break connectors. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#42-incremental-polling-jdbcodbc)

## 4.3 Log-based CDC

- Read the transaction log (MySQL binlog, PostgreSQL WAL, MongoDB oplog/change streams, SQL Server CDC). Debezium is the dominant open-source platform. Advantages: captures inserts/updates/deletes, every intermediate state, minimal source load, sub-second latency, stable per-row ordering. Typical pipeline: initial snapshot, then stream the log from the snapshot's LSN/position; persist resume tokens/offsets for recovery. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#43-log-based-cdc)

## 5.1 Kafka, Kinesis, Pub/Sub

- Kafka (MSK, Confluent, Redpanda): de facto standard, open protocol, strongest ecosystem, highest throughput. Best for multi-cloud and complex stream processing. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#51-kafka-kinesis-pubsub)
- Kinesis Data Streams: AWS-native, shard-based; 2026 trend favors MSK for new AWS deployments unless small/serverless. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#51-kafka-kinesis-pubsub)
- Pub/Sub: GCP-native, serverless, regional exactly-once (2024). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#51-kafka-kinesis-pubsub)

## 5.2 Exactly-once semantics

- At-most-once / at-least-once / exactly-once. Kafka: idempotent producers + transactions API. Kinesis: KCL checkpoints + idempotent downstream. Pub/Sub: regional exactly-once API. Practical guidance: make consumers idempotent, default to at-least-once, invoke exactly-once only when duplicates are costlier than the coordination. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#52-exactly-once-semantics)

## 5.3 Order, partitioning, back-pressure

- Partition key sets parallelism and ordering (same key → same partition → in-order). Hot partitions are the primary failure mode. Handle back-pressure at the consumer: buffer (memory), drop (loss), or pause the source (propagation). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#53-order-partitioning-back-pressure)

## 6.1 ETL vs ELT

- ETL: transform before load (Informatica, Talend, SSIS, Glue). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#61-etl-vs-elt)
- ELT: load raw, transform in-warehouse (Snowflake/BigQuery/Redshift/Databricks). Default 2026 pattern; keeps raw history, decouples ingest from transformation. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#61-etl-vs-elt)

## 6.2 The modern data stack

- Ingest/EL: Fivetran, Airbyte, Meltano (Singer), Stitch. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Transform/T: dbt (needs an orchestrator; doesn't extract/load). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Orchestration: Airflow, Dagster, Prefect, Mage. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Warehouse: Snowflake, BigQuery, Databricks, Redshift, ClickHouse, MotherDuck. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Reverse ETL: Hightouch, Census. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Catalog/governance: DataHub, OpenMetadata, Atlan, Collibra. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)
- Observability: Monte Carlo, Bigeye, Lightup, Soda. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#62-the-modern-data-stack)

## 6.3 Selection guidance

- Small all-SaaS team: Fivetran + dbt Cloud + Snowflake + Hightouch (watch Fivetran per-MAR pricing). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#63-selection-guidance)
- Mid-size open-source: Airbyte + dbt Core + Snowflake/BigQuery + Airflow. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#63-selection-guidance)
- Code-first team: Meltano + dbt Core + Airflow + Snowflake/BigQuery. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#63-selection-guidance)

## 7.1 Sampling frame

- The operational list of population members that can actually be reached - rarely identical to the target population. This gap creates coverage bias. Document the frame explicitly at design time; if it doesn't match the population, no sample size or design can fix the resulting bias. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#71-sampling-frame)

## 7.2 Response bias

- Bias is systematic distortion; it does not shrink with sample size. Forms: nonresponse, acquiescence (yea-saying), social desirability, recall, order effects, mode effects, selection bias. Mitigations: track response rate, reverse-coded items, anonymous administration, randomize order, demographic benchmarking. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#72-response-bias)

## 7.3 Practical design

- Pilot with 10-20 respondents; use vertical scales for mobile; cap length (completion falls past 5-7 min); use attention checks sparingly; pre-register the analysis plan for high-stakes work. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#73-practical-design)

## 8.1 Probability sampling

- Every unit has a known, non-zero selection probability - the only basis for valid frequentist inference. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#81-probability-sampling)
  - SRS: equal probability 1/N; the reference design. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#81-probability-sampling)
  - Stratified: sample within mutually exclusive strata. Proportional allocation (size-proportional) vs Neyman optimal allocation (size × stdev; minimizes variance for fixed n). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#81-probability-sampling)
  - Cluster: randomly select clusters, then sample within. Loses precision vs SRS (design effect; effective sample size n / DEFF). — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#81-probability-sampling)
  - Systematic: every kth element after a random start; biased if the frame has periodicity matching k. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#81-probability-sampling)

## 8.2 Non-probability sampling

- Selection probability unknown or zero; treat as exploratory unless you can model selection. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
  - Convenience: whoever is at hand. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
  - Quota: hit target subgroup counts; biased on non-quota dimensions. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
  - Snowball: respondents refer others; good for hidden populations. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
  - Purposive/judgment: researcher selects informative units; fine for case studies, never for population estimates. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
- Modern hybrid: online panel + post-stratification weighting - weight non-probability panel responses to population marginals. Reduces but does not eliminate selection bias on outcome-correlated dimensions not in the weighting variables. — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)
- <!-- cross-hub-map --> — [source](https://llms-explorer.com/sources/mdb-context-hub/da-3-data-acquisition-sampling/#82-non-probability-sampling)

## Project ideas

- Build a resumable API ingestion client that persists its pagination cursor after every page, so a partial failure mid-crawl can resume instead of restarting from page one. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Set up log-based CDC (via Debezium against a database's transaction log) for a table where hard deletes and backdated updates matter, instead of relying on incremental polling that would miss them. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Design a stratified sampling plan with Neyman optimal allocation for a survey where subgroup variances differ meaningfully, instead of defaulting to simple random sampling. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Stand up a small modern-data-stack ingestion pipeline — Airbyte or Fivetran into a warehouse, transformed with dbt, orchestrated with Airflow or Dagster — and choose ELT over ETL as the default load pattern. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Where this helps

- Any project where just pulling the data would skip past a real coverage gap in the sampling frame or a schema drift that only surfaces once analysis is already underway. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Choosing between REST, GraphQL, and gRPC for a new data-acquisition integration, based on whether over/under-fetching, caching, or internal service-to-service throughput matters most. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Deciding between bulk export, incremental polling, and log-based CDC for keeping a downstream copy of a source table current, based on how much you need to capture deletes and intermediate states. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Designing a survey or panel where response bias — nonresponse, social desirability, acquiescence — is a real risk that larger sample size alone will not fix. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Antipatterns

- Treating data acquisition as a pure logistics problem — just pull the data — without checking upfront whether the population, sample frame, or schema will actually support the intended analysis. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Using bulk SELECT * exports as the ongoing sync mechanism for a large, actively-changing operational table, instead of incremental polling or log-based CDC. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Scraping a site that has no public API without checking robots.txt, throttling, and the CFAA/Terms-of-Service risk surface first. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Treating a non-probability sample — convenience, quota, or purposive — as if it licensed a population-level statistical estimate, when only probability sampling designs support that kind of inference. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- Incremental polling based on an updated_at column or auto-increment ID structurally cannot see hard deletes or backdated updates — those require log-based CDC to capture correctly. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Post-stratification weighting on a non-probability online panel reduces but does not eliminate selection bias on any outcome-correlated dimension that isn't included in the weighting scheme. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Cluster sampling loses statistical precision relative to simple random sampling of the same size, the design effect, so a naive sample-size calculation that ignores clustering will understate the true margin of error. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Web scraping's legal footing after hiQ v. LinkedIn covers only the CFAA angle for publicly accessible data — Terms of Service civil claims, copyright, and site-specific access-control bypass remain separate legal risks. — [source](https://llms-explorer.com/tree/data-acquisition-and-sampling/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [Data Acquisition and Sampling](https://llms-explorer.com/downloads/sources/mdb-context-hub/da-3-data-acquisition-sampling.md)
