{"name":"ScaleDown","description":"Small language models that each do one job: compress long context, summarize, pull structured data out of text, and sort text into categories you define. A fraction of a general-purpose model's cost, and faster. Prepaid credits bought over a payment rail; no account, no API key.","docs":{"agent_card":"https://agents.scaledown.ai/.well-known/agent-card.json","llms":"https://agents.scaledown.ai/llms.txt","mpp":"https://agents.scaledown.ai/.well-known/mpp.json","openapi":"https://agents.scaledown.ai/openapi.json","skill_md":"https://agents.scaledown.ai/skill.md"},"endpoints":{"balance":"GET /credits: read the prepaid token balance behind your operator token, and what it is worth at the current rate. Free.","topup":"POST /credits: whole-dollar top-ups (optional `usd` 1-50, default 5). Crypto rails (Tempo, Solana, x402 Base) settle FLAT from $1 with no fee. Card/Link is $5 multiples and adds a $0.20 service fee per $5 ($5 bills $5.20, $10 bills $10.40); the fee is a fee, so the credit is the same either way. $1 credits 20,000,000 tokens.","extract":"POST /extract: Turn unstructured text into structured data: describe each field you want in plain English, and get back the values found in the text, each with a confidence score and the exact position it came from. Metered against your prepaid balance at $0.05 per 1M tokens, flat, with no tiers.","classify":"POST /classify: Sort text into categories you define: write a yes/no question describing each category, and get back a score for every one plus the winning label, with nothing to train and no example data to collect. Metered against your prepaid balance at $0.05 per 1M tokens, flat, with no tiers.","summarize":"POST /summarize: Condense long text into a short, readable write-up in the model's own words, shaped by an instruction you give it: bullet points, one sentence, a word limit, a language, or a topic to keep to. Metered against your prepaid balance at $0.05 per 1M tokens, flat, with no tiers.","compress":"POST /compress: Strip the padding out of a long context before another model has to read it: send the context together with the question it exists to answer, and get back a much shorter version that keeps whatever bears on that question. Metered against your prepaid balance at $0.05 per 1M tokens, flat, with no tiers."},"audience":"agents","supported_rails":["tempo_mpp","x402_base","solana_mpp","stripe"],"capabilities":{"extract":{"endpoint":"POST /extract","request_example":{"entities":{"auditor":"The audit firm, if named."},"text":"Acme completed a SOC 2 Type II audit in 2026, performed by Example Assurance LLP."},"response_example":{"entities":[{"confidence":1,"context":"Acme completed a SOC 2 Type II audit in 2026, performed by Example Assurance LLP.","end":80,"start":59,"text":"Example Assurance LLP","type":"auditor"}],"metering":{"balance_remaining_tokens":19999847,"charged_tokens":153,"charged_usd_at_current_rate":"0.000008","endpoint":"extract","tokens_source":"measured"}},"summary":"Turn unstructured text into structured data: describe each field you want in plain English, and get back the values found in the text, each with a confidence score and the exact position it came from.","use_cases":["Contract analysis: parties, effective dates, termination clauses, governing law","Resume parsing: name, email, skills, employers, job titles, education","Medical records: diagnoses, medications, dosages, lab values, dates of service","Financial filings: revenue figures, dates, named parties, regulatory references","Support tickets: product names, error codes, account identifiers","News and media monitoring: people, organizations, locations, quoted statements","Product listings: names, prices, product codes, brands, specifications"],"what_it_does":"Ordinary entity extractors only recognize a fixed list of things: people, places, organizations. This one reads the descriptions you write, so it can pull contract parties, invoice totals, part numbers, dosages, error codes, or anything else you can describe in a sentence. Every result carries a confidence score you can filter on, the character offsets where it was found, and the surrounding text, so a value can be checked against its source or sent to a person for review. There is nothing to train and no examples to label: rewriting the description is how you change the behavior. It replaces regular expressions and hand-built parsing pipelines, and it is not limited to flat fields: an entity can be a nested object or a repeated list, which comes back assembled under `structured_result`."},"classify":{"endpoint":"POST /classify","request_example":{"labels":[{"name":"critical","rubric":"Does this describe a service fully down or data loss right now?"},{"name":"low","rubric":"Does this describe a cosmetic issue or a question?"}],"text":"The server has been down for three hours and customers cannot check out."},"response_example":{"labels":[{"label":"critical","rubric":"Does this describe a service fully down or data loss right now?","score":0.9999852612119633},{"label":"low","rubric":"Does this describe a cosmetic issue or a question?","score":0.000014738788036782268}],"reasoning":"The text explicitly states the server is down and customers cannot check out, indicating a fully down service.","scores":{"critical":0.9999852612119633,"low":0.000014738788036782268},"top_label":"critical","metering":{"balance_remaining_tokens":19999847,"charged_tokens":153,"charged_usd_at_current_rate":"0.000008","endpoint":"classify","tokens_source":"measured"}},"summary":"Sort text into categories you define: write a yes/no question describing each category, and get back a score for every one plus the winning label, with nothing to train and no example data to collect.","use_cases":["Support ticket triage: route to the billing, technical, or account queue","Content moderation: score against a spam, abuse, or policy rubric and act on a threshold","Intent detection: question, complaint, feedback, or cancellation","Topic tagging: label documents by subject for downstream routing or filtering","Lead qualification: score inbound inquiries on urgency or deal size","Inbox sorting: file incoming mail into categories without a rules engine"],"what_it_does":"You supply the categories, and for each one a rubric: a yes/no question phrased so that \"yes\" means the category applies. The model scores the text against every question and normalizes the results into a set of scores that add up to 1, returning the strongest as `top_label`, the full score map, and a short written reason for the answer. The scores are relative rather than absolute confidence: they say how the text divides between the categories you gave, so a top score of 0.85 means 85% of the mass went there rather than 85% certainty, and a threshold is yours to apply on the `scores` map. The winning label is stable across runs; the exact scores move slightly. Changing your categories means rewriting a sentence, not collecting data and retraining a model."},"summarize":{"endpoint":"POST /summarize","request_example":{"instructions":"One sentence, plain language.","text":"Acme Corp announced quarterly results today. Revenue rose 14 percent to 4.2 billion dollars, driven by cloud services."},"response_example":{"input_chars":118,"latency_ms":192,"output_chars":91,"summary":"Acme Corp reported a 14 percent revenue increase to $4.2 billion, driven by cloud services.","metering":{"balance_remaining_tokens":19999847,"charged_tokens":153,"charged_usd_at_current_rate":"0.000008","endpoint":"summarize","tokens_source":"measured"}},"summary":"Condense long text into a short, readable write-up in the model's own words, shaped by an instruction you give it: bullet points, one sentence, a word limit, a language, or a topic to keep to.","use_cases":["Document review: contracts, research papers, policy documents","News digests: condense articles into a paragraph or a few bullets","Customer feedback: summarize support tickets, reviews, or survey responses at scale","Meeting transcripts: turn a call recording into the decisions and action items","Financial reporting: the key figures and decisions out of an earnings call or a filing","Pipeline pre-processing: shorten retrieved passages before handing them to a larger model"],"what_it_does":"It rewrites the text rather than stitching together sentences lifted out of it, so the result reads as prose instead of a set of clipped quotes. The sampling settings are fixed to a configuration tuned for staying faithful to the source, which keeps output stable and low-noise across a pipeline instead of varying run to run. The `instructions` field steers shape, length, language, and focus without loosening that. Built for long inputs: documents, call transcripts, and long message threads."},"compress":{"endpoint":"POST /compress","request_example":{"context":"The quarterly earnings report for FY2026 shows revenue of $4.2 billion. We also saw strong performance across all regional segments globally. Weather conditions in Q3 were noted as a minor operational factor. The board met twice to review internal governance documentation. Net income increased 18 percent year over year, driven by margin expansion. Office supply procurement was streamlined across all twelve locations. Customer retention reached an all-time high of 94.7 percent this year.","prompt":"What drove the increase in net income?","scaledown":{"rate":"auto"}},"response_example":{"latency_ms":49,"results":{"compressed_prompt":"The quarterly earnings report for FY2026 shows revenue of $4.2 billion. We also saw strong performance across all regional segments globally. Weather conditions in Q3 were noted as a minor operational factor. The board met twice to review internal governance documentation. Net income increased 18 percent year over year, driven by margin expansion. Office supply procurement was streamlined across all twelve locations.","compressed_prompt_tokens":76,"compression_ratio":0.6333333333333333,"original_prompt":"What drove the increase in net income?","original_prompt_tokens":120,"success":true},"metering":{"balance_remaining_tokens":19999847,"charged_tokens":153,"charged_usd_at_current_rate":"0.000008","endpoint":"compress","tokens_source":"measured"}},"summary":"Strip the padding out of a long context before another model has to read it: send the context together with the question it exists to answer, and get back a much shorter version that keeps whatever bears on that question.","use_cases":["Retrieval pipelines: shrink fetched passages before they go into the prompt","Question answering over a long document that would not otherwise fit","Long conversations: compress the history instead of dropping the start of it","Code review: large files or diffs passed along as context","Repeated pipelines: the same compression step run across many documents"],"what_it_does":"The compression knows what the question is, so it keeps the parts that bear on it and drops the parts that do not, rather than trimming blindly from one end. Your question passes through intact; the surrounding context is what shrinks. What comes back is a drop-in replacement for the context and prompt you were about to send, so neither the model you call nor your own code has to change. Reach for it to cut the bill on a long call, to fit a document that would otherwise be too big, to shorten the wait for the first token back, or to make a job practical on a smaller and cheaper model."}},"identity":"Your AgentScore passport token names the account the prepaid balance belongs to. Nothing else is required and no personal information is collected. A call with no token returns 401 with a verify_url (a one-click account sign-in, no identity documents) and a poll_url that returns the operator token, so any HTTP client can bootstrap; the AgentScore pay CLI does it for you.","limits":["A single call is capped at 250,000 characters, roughly 62,500 tokens.","Text input only. Image and PDF input (read by optical character recognition) and the asynchronous batch interface are not available.","A per-buyer rate limit applies. Hitting it returns 429 and never touches your balance."],"pricing":{"metering":"You are billed for the tokens the model counted for your input, and only for calls that succeed. Your balance holds tokens rather than dollars, bought at the rate listed here, so a later price change never touches credit you already own.","rate":"$0.05 per 1M tokens, flat, with no tiers","topup":"whole-dollar top-ups (optional `usd` 1-50, default 5). Crypto rails (Tempo, Solana, x402 Base) settle FLAT from $1 with no fee. Card/Link is $5 multiples and adds a $0.20 service fee per $5 ($5 bills $5.20, $10 bills $10.40); the fee is a fee, so the credit is the same either way. $1 credits 20,000,000 tokens"},"tagline":"Four purpose-built text models, prepaid by the token and payable by an agent.","why_scaledown":["Each model is trained for one task, so a job that would otherwise go to a general-purpose large model costs a fraction as much and returns faster.","You describe what you want in plain English: the fields to pull, the categories to sort into, the shape of the summary. There is nothing to fine-tune, no training data to collect, and no examples to label.","Extraction returns a confidence score, the character offsets, and the surrounding text for every value, so a result can be checked against its source or sent to a person for review.","One prepaid balance covers all four endpoints, and it belongs to your AgentScore account rather than to a key, so rotating your token never strands it.","Nothing you send is kept. Token counts and timestamps are recorded so a balance can be metered; the text itself is never written down."]}