AI HTS Classification: Where It Helps, Where It Stops
AI HTS classification can speed up tariff coding, but it cannot carry the legal duty. Where automation helps, where it fails and where people must review.
AI HTS classification is no longer experimental. Language models can read a product description, search the Harmonized Tariff Schedule, compare candidate codes and draft reasoning in seconds, which is why importers and brokers are adding them to their workflows. Used well, they shorten the slowest part of classification. Used as an oracle, they produce confident codes that nobody can defend when CBP asks why.
This article is for importers, customs brokers and e-commerce teams deciding how far to rely on automation. It sets out what AI does well, where it fails, and where a qualified person must stay in the loop. HTS Pilot is one such tool, so we are not neutral on whether automation is useful; we have tried to be specific about its limits, including our own.
Key takeaways
- The importer of record must use reasonable care to declare classification. That responsibility does not move to software.
- AI is strong at reading descriptions, searching the schedule, retrieving rulings and drafting reasons at scale.
- AI is weak where classification depends on facts not in the description, on judgement calls such as essential character, or on recent changes it does not know.
- A trustworthy tool shows its grounds, uses only real codes from the current tariff version, and routes uncertain cases to people.
- People should review composite goods, sets, parts, close calls and high-value lines, and keep the decision on file.
The legal frame: who decides
Under 19 U.S.C. 1484(a)(1), the importer of record, "using reasonable care", files the declared value, classification and rate of duty. CBP can accept the classification or change it, and for certainty before importing, an importer can request a binding ruling under 19 CFR Part 177. Nothing in that framework recognises a software output as a decision.
So the useful question is not whether AI can replace classification work, but whether it helps a person reach a defensible code faster and with better evidence. Our customs misclassification penalties guide explains what is at stake when the code is wrong.
Where AI helps
Reading messy descriptions. Supplier descriptions are inconsistent, abbreviated and sometimes in other languages. Models are good at extracting materials, percentages, construction, function and user from free text.
Searching a large schedule. By our count of USITC's export of Revision 20, the HTS has about 20,000 10-digit statistical lines across chapters 1 to 99. Combining keyword search, semantic search and the schedule's tree structure finds candidates a person might miss.
Retrieving precedents. CBP rulings in CROSS are the best evidence of how CBP reads the schedule. Retrieval systems can bring the closest rulings to the reviewer's screen, which is slow to do by hand. See CBP binding rulings for how to judge them.
Applying rules consistently. Some legal notes can be encoded as hard rules. Chapter 61 note 1, for example, limits the chapter to knitted or crocheted articles, so a woven garment can be excluded from chapter 61 before any model reasons about it.
Asking for missing facts. A system can see that the candidate codes differ on one attribute, knitted or woven, cotton or polyester, and ask that question instead of guessing.
Working at scale. For a catalogue of thousands of SKUs, automation turns weeks of first drafts into hours, leaving people to review the hard cases.
Where AI stops
Facts that are not in the description. The code often depends on facts the text does not state. In ruling N353266, CBP sent a water bottle to its laboratory to confirm it was a vacuum vessel before classifying it in heading 9617. No model can infer a vacuum, a fibre percentage or a fabric weight that nobody wrote down.
Judgement calls. GRI 3(b) asks which material or component gives a composite good its essential character. That judgement weighs role, bulk, weight and value differently for each product. Models can argue either side convincingly, which is exactly why a person should decide.
Recent changes. The HTS was revised 32 times in 2025 and 20 times in 2026 by late September, and CBP revokes rulings it no longer agrees with. A model that relies on what it learned in training rather than the current schedule and current rulings will be out of date.
Plausible but wrong reasoning. Validation can stop a system from inventing a code or a source. It cannot guarantee the reasoning about a real code is correct. Confident prose is not evidence.
Duty is more than the code. Chapter 99 additional duties, AD/CVD orders and agency requirements depend on origin and on measures outside the 8-digit rate.
What a responsible AI classification tool should do
| Requirement | Why it matters |
|---|---|
| Uses only declarable codes from one current tariff version and one market | Prevents invented or expired codes and mixed schedules |
| Shows alternatives and why each was chosen or ruled out | Lets a reviewer check the reasoning, not just the answer |
| Cites official sources and rulings, with access dates | Gives the importer evidence for its file |
| Applies legal notes and GRIs as explicit rules where possible | Keeps hard exclusions out of model guesswork |
| Reports status and confidence honestly | Separates "proposed" from "needs more info" and "needs expert review" |
| Asks for missing deciding facts | Avoids guessing between knitted and woven, or cotton and polyester |
| Keeps an audit trail of who decided what, on which version | Supports reasonable care after the fact |
Questions to ask before you adopt a tool
Whatever product you evaluate, including ours, these questions separate useful automation from a black box:
- Which tariff data does it use, and how current is it? Ask how quickly a new USITC revision reaches the tool, and whether each result records the version used.
- Can it return a code that does not exist? A tool should check every proposed code against the declarable lines of the current schedule.
- What does it show besides the code? Look for alternatives, reasons for and against each, and sources a reviewer can open.
- What happens when the description is incomplete? It should ask for the missing fact or mark the result as needing more information, not fill the gap with a guess.
- How does it handle rulings? Ask whether revoked rulings are excluded and whether cited rulings are linked.
- What does it do with uncertain cases? There should be a status for "needs review" and a place where a person makes the call.
- Does it claim an accuracy figure? If so, ask what was measured, on which products and against whose answers. A headline percentage says little about your catalogue.
- What is logged? You need to reconstruct, months later, who accepted which code and on what grounds.
Test with your own products, especially the difficult ones, before relying on any tool for live entries.
Where the human review point belongs
Review is not a formality at the end. It belongs wherever the evidence is thin or the rules require judgement:
- Composite goods, sets, parts and unfinished goods, where GRI 2 and GRI 3 apply. See General Rules of Interpretation.
- Missing deciding facts, until the description is completed.
- Close candidates, when two codes are nearly equally supported.
- Divergent duties, when plausible codes lead to very different duty.
- High-value or repeated imports, where a binding ruling may be worth requesting.
- Changes, when a new revision or a revoked ruling affects a code already in use.
A workable pattern: the tool drafts, the reviewer checks reasons and sources, decides, and the decision is logged with the tariff version. Over time, reviewer decisions show where the tool is reliable and where it is not. Our guide to the tariff classification audit trail covers what that record should contain.
How HTS Pilot applies this
HTS Pilot was built around these limits. It extracts attributes from the description with rules and a language model, keeping a model value only when it quotes the input text; narrows the schedule by chapter and heading; retrieves CBP rulings as precedents; and removes codes that the GRIs and legal notes rule out before the model compares the rest. The model may only choose among candidate codes and cite sources it was given, and fixed checks confirm the code is a declarable line in the right market and version. Weak grounds produce "Needs more info" or "Needs expert review", and multi-material goods, sets, parts and unfinished goods go to a reviewer. Reviewer decisions are stored and used to retrain its ranking. Its results are suggestions for reference, not official classification decisions.
Summary
AI HTS classification is useful for the parts of the job that are search, extraction and drafting, and unreliable for the parts that are judgement, unstated facts and fresh legal change. Use it to reach a well-evidenced draft quickly, keep a qualified person at the review points that matter, and keep the reasoning on file. The method behind it is the same one described in how to find HTS code numbers.
Frequently asked questions
Can AI classify HTS codes accurately?
AI can propose plausible codes quickly and handle large catalogues, but accuracy depends heavily on the product description and on how the system is built. A language model can reason wrongly about a real code. Use AI output as a draft with its reasons and sources, have a qualified person review uncertain cases, and keep the decision and its grounds on file.
Who is responsible if an AI tool gives the wrong code?
The importer of record. Under 19 U.S.C. 1484, the importer must use reasonable care to declare the classification and rate of duty. Using a tool does not transfer that duty. A tool can support reasonable care by showing its grounds and sources, but the importer must still be able to explain why the declared code is right.
What should an AI classification tool show me?
At minimum: the proposed 10-digit code, alternatives it considered, why each was chosen or ruled out, the official sources and rulings it relied on, the tariff version used, and a confidence or status that tells you when to review. A tool that shows only a code and no reasons gives you nothing to check and nothing to keep on file.
When should a person review an AI-suggested code?
Always for multi-material goods, sets, parts and unfinished goods, where GRI 2 and 3 apply; when key facts such as fibre content or construction are missing; when two candidate codes are close; when candidate codes carry very different duties; and for high-value or repeated imports. For those, also consider a binding ruling request to CBP.
Sources
The official texts and pages this article relies on. Check them for the current version before you act.
- 19 U.S.C. 1484, Entry of merchandise (reasonable care) - govinfo govinfo.gov
- 19 CFR Part 177, Administrative Rulings - govinfo govinfo.gov
- General Rules of Interpretation, HTS 2026 Revision 20 - USITC hts.usitc.gov
- FAQs about tariff classification and the HTS - U.S. International Trade Commission usitc.gov
- CBP ruling N353266: water bottle classified after laboratory analysis rulings.cbp.gov
This article is general information, not legal advice and not a classification decision. Tariff texts, rates and rulings change: check the current official sources, and ask the customs authority for a binding ruling where the answer matters.