Payroll & Tax

How to study marketplace-facilitator sales-tax evidence

A reproducible study protocol for whether marketplace-collected tax is separated from seller-collected tax using jurisdiction, order, facilitator, settlement, filing, and ledger evidence, with frozen populations, evidence coverage, decision rights, and limitations.

A reproducible study protocol for whether marketplace-collected tax is separated from seller-collected tax using jurisdiction, order, facilitator, settlement, filing, and ledger evidence, with frozen populations, evidence coverage, decision rights, and limitations.

Key takeaways

  • Freeze the population, cutoff, clock, and status definitions before calculating results.
  • Report missing evidence, returned work, exclusions, and reopened items beside the headline measure.
  • Keep preparation with the bookkeeping team and judgment, approval, release, and policy decisions with the authorized client owner.

Research question and preregistered definitions

Ask one narrow question: for a frozen population, how many records meet the event, what states explain the remainder, and how much evidence is missing? Before extraction, record the numerator, denominator, observation period, local time zone, cutoff, eligible states, exclusions, reopen rule, pause rule, and treatment of late-arriving records. A percentage without its underlying counts is not decision-grade.

The denominator is every eligible one marketplace order tax component for one seller entity and destination jurisdiction in consecutive periods. The primary numerator is every record meeting this event: a tax component whose collection responsibility or ledger disposition is unsupported or inconsistent across marketplace and bookkeeping records. Retain open, excluded, and indeterminate records in separate tables with reasons. Never remove a record because its support is inconvenient or because it arrived late. The authorized finance owner should approve eligibility and exception rules before the pilot.

Use system timestamps where they are fit for purpose. State whether elapsed time means continuous clock time or agreed working time. Preserve local time and UTC when teams cross time zones. If an item reopens, either treat the first closure as provisional or create a new episode; choose once, before seeing results. Report reopened counts because apparently quick closure can conceal repeated returns.

Population and data collection

Select one workflow, one entity or an explicitly listed entity group, and consecutive periods. Avoid a handpicked clean week. Each study row should contain order token, seller entity, marketplace, destination jurisdiction, order date, taxable base, tax amount, collector indicator, settlement reference, ledger mapping, filing treatment, reviewer, and evidence link. Stable identifiers matter: updates to an existing item must not create a second apparent observation, while two genuinely separate events must not be collapsed because their amounts happen to match.

Preserve each raw export, extraction timestamp, report parameters, schema version, row count, control total, and file hash where practical. Reconcile the extract to an independent system report when one exists. Log filters, inaccessible systems, manual supplements, duplicate identifiers, blank timestamps, and post-extraction additions. A larger dataset cannot cure a broken lineage, so evidence coverage belongs beside the primary result.

Collect the least sensitive data the question needs. Replace names with stable study identifiers when identity is irrelevant. Exclude bank credentials, complete account numbers, tax identifiers, compensation detail not needed for the test, and unrelated free text. Store any reidentification key separately, restrict access by role, and follow the client's approved retention and deletion schedule.

Calculations and reporting

Report the primary event count divided by the frozen eligible population, with the numerator and denominator printed beside the percentage. Also show open, returned, excluded, indeterminate, missing-evidence, late-arriving, and reopened counts. For elapsed time, show a median and useful age bands, plus the oldest open items. An average alone can hide a small group of very old records.

Break down the result by marketplace, entity, jurisdiction, product tax class, collector state, filing treatment, amount band, and exception reason. Suppress or combine small cells where needed to protect people and counterparties. A difference between groups is descriptive. It does not establish that a person, staffing model, location, or application caused the result. Volume, complexity, policy changes, migrations, outages, source delays, reviewer capacity, and changes in evidence quality are plausible confounders.

Publish a population reconciliation, data-quality table, state counts, age distribution, exception table, reviewer-disagreement table, and change log. Pair every chart with counts. Keep historical extracts immutable and issue corrections through versioned copies. A reviewer should be able to reproduce the total and understand why a later version differs.

Interpretation for offshore bookkeeping

Use the findings to improve instructions, access, evidence flow, and escalation, not to manufacture a market benchmark. Review actual exceptions before changing headcount or deadlines. When the measure moves, first test whether the population, cutoff, system, rule, evidence coverage, reviewer assignment, or approval path changed. Only then consider an operational explanation. One favorable period is not proof that a control is effective.

For a Philippines-based support team, document overlap hours, handoff cutoff, relevant holidays, source-system availability, named escalation route, and maximum waiting time for unresolved items. Use named accounts, multifactor authentication, and least-privilege access. Where the client's risk assessment requires separation, keep source maintenance, preparation, accounting approval, payment release, and period locking with distinct authorized roles.

The niche-specific conclusion is that whether marketplace-collected tax is separated from seller-collected tax using jurisdiction, order, facilitator, settlement, filing, and ledger evidence can be evaluated only when the client owns definitions and decision rights while the bookkeeping team owns orderly preparation and escalation. That boundary lets an offshore bookkeeper add capacity without quietly inheriting authority reserved for management.

Order-channel responsibility test

Construct the population from order-level marketplace exports and direct-channel records, not merely from ledger tax accounts. For every selected order, retain destination, product tax class, taxable base, tax charged, collector indicator, refund history, marketplace identity, and settlement reference. Reconcile gross order tax to marketplace reports before comparing net deposits, because settlements may net commissions, refunds, reserves, and other deductions.

Responsibility can vary by jurisdiction, period, channel, and transaction facts. The study therefore records the evidence used by the client's tax owner instead of creating a universal facilitator rule. A bookkeeper may apply an approved jurisdiction table and flag conflicts; they should not decide nexus, registration, exemption validity, product taxability, or filing positions. Marketplace labels are inputs, not conclusive legal evidence.

Filing-to-ledger bridge

Bridge seller-collected tax, facilitator-collected tax, refunds, adjustments, remittances, and ending liabilities separately. Test orders around registration changes, month end, destination changes, and amended returns. Mixed baskets and partial refunds deserve their own examples because allocating tax by gross order can conceal errors. Report unsupported collector indicators and unreconciled settlement differences as separate exception families. This design identifies where evidence breaks without claiming that an observed difference is tax due.

Limitations and uncertainty

This brief contains no private dataset, prevalence estimate, market benchmark, causal effect, savings claim, or provider comparison. A pilot describes only the selected population under its declared rules. Small populations produce unstable rates. Missing timestamps may be systematic rather than random. Different systems can record the same business event at different stages, and decisions made outside the system may be absent.

Comparisons across teams or periods require equivalent definitions, populations, clocks, systems, and evidence coverage. Even then, treat a difference as a prompt for review. Do not rank employees, infer misconduct, promise a financial outcome, or claim control effectiveness from this measure alone. Qualified accounting, audit, tax, legal, security, payroll, treasury, and statistical owners should review issues within their remit.

Before reuse, disclose the sample size, period, entities, exclusions, missingness, system changes, codebook revisions, reviewer disagreement, conflicts, and tolerance choices. Archive the protocol with results. An honest limitation and traceable denominator are more useful than a precise-looking percentage that cannot be reconstructed.

Implementation checklist

  1. Name the workflow and decision owners. 2. Freeze the population and period. 3. Approve the event, exclusion, reopen, clock, and pause rules. 4. Export and reconcile the population. 5. Minimize sensitive fields. 6. Apply the codebook. 7. Independently recode a sample. 8. Publish counts, distributions, missingness, and exceptions. 9. Review source records before changing the workflow. 10. Version the protocol and retain evidence under the approved schedule.

Run consecutive periods long enough to observe ordinary variation, while repairing clear access or evidence defects immediately. Keep protocol defects separate from operational findings. A useful pilot ends with clearer owners, fields, stops, and escalation even when the headline measure remains uncertain.

Sources and checked dates

Evidence map

These notes connect bounded statements on this page to the listed public sources. They do not turn operational interpretations into empirical findings.

  1. GAO presents principles for designing, implementing, and operating internal control and evaluating deficiencies; this article applies those principles by analogy and does not claim GAO prescribes the proposed bookkeeping metric.
  2. PCAOB AS 1105 and AS 1215 address evidence and documentation in PCAOB audit contexts; they support the distinction between a recorded conclusion and inspectable support, but do not turn this protocol into an audit procedure.
  3. NIST materials support explicit access roles and protection of data integrity; they do not decide accounting treatment, payment authority, or staffing performance.
  4. IRS recordkeeping guidance explains that supporting records should substantiate business transactions; retention and tax treatment still depend on the actual record and applicable requirements.

Turn the protocol into a reviewable handoff

Define the source records, preparation fields, stops, reviewer, and escalation route before assigning the workflow.

Discuss a controlled bookkeeping scope

Listed sources

  1. U.S. GAO, 2025 Green Book
  2. PCAOB, AS 1105: Audit Evidence
  3. PCAOB, AS 1215: Audit Documentation
  4. PCAOB, AS 2201: An Audit of Internal Control Over Financial Reporting
  5. NIST, Cybersecurity Framework 2.0
  6. NIST, Data Integrity
  7. NIST, Role Based Access Control
  8. IRS, Recordkeeping

Related research