Skip to content

Forge shapes AI models around your work.

Create custom and fine-tuned models around your business’s tasks, examples, and requirements. Start with data you already have, or with the task itself. Traces are optional. The result is evaluated against the work it needs to do.

What Forge does for your team

Forge is Sparset’s model-development offering. It takes a defined job, the examples and requirements that describe it, and optional traces, and produces a custom or fine-tuned model evaluated against that job. Training is the method, not the outcome: a candidate ships when it meets the agreed criteria, or the work stops with a record of why it did not. Forge develops the model. Stack or Edge can operate it.

How you can begin

Three starting points that belong in a Forge conversation. Traces are optional. A project still needs representative evidence of the task, and a way to assess results.

Bring your examples

You have approved examples, labels, corrections, or a task dataset. We assess suitability, prepare the data, and build an evaluation baseline.

Use selected traces

You have relevant interaction records, or you choose to instrument a workflow. Trace collection is optional. We agree the permitted fields and use, then convert approved material into examples.

Start with the task

You know the problem but do not yet have a usable dataset. We assess feasibility and scope a representative example-building effort with you. No traces is a valid starting point; no evidence at all is not a promise.

From a defined job to a candidate you can accept

Forge runs the development process. Your team defines the job, authorizes the data, and accepts or rejects each candidate.

  • Define

    Agree the job, the failures, and what success means. Your team explains the workflow and the constraints.

  • Prepare

    Turn approved material into training and evaluation examples, kept separate. Your team supplies domain judgment and reviews what may be used.

  • Adapt

    Choose a supported base model and adaptation method, and train a candidate. Your team resolves task-specific questions; what we can support is confirmed during evaluation.

  • Evaluate

    Compare the candidate with the baseline on agreed tasks, including failures. Your team reviews the evidence and decides whether to accept, iterate, or stop.

  • Deliver

    Hand over the agreed artifact, the evaluation summary, and the intended-use notes. Your team accepts the deliverable and decides the serving path; a new candidate does not replace production on its own.

How you know a candidate is better

See what improved, and what still needs work. The comparison is against the agreed baseline, not a published score.

Compare with the baseline

The candidate is judged against your current approach on agreed tasks and criteria, including cases it was not trained on. A favorable result is an evaluation outcome, not an automatic consequence of fine-tuning.

Failures and tradeoffs

A model can improve on the target job and get worse in other ways. The report includes representative failures so the acceptance criteria can make those tradeoffs explicit.

Serving suitability

Quality is considered alongside the intended latency, throughput, and infrastructure. Compatibility with the serving environment is confirmed for the project, not assumed.

Shape the model with Forge. Operate it with Stack or Edge.

Forge is the model-development offering. Stack and Edge operate inference. You can discuss Forge before choosing a serving environment; bundling is agreed per engagement, not assumed.

  • Stack

    For engineering and platform teams that want the model operated in production: planning, deployment, and the operating loop on GPU capacity you own, rent, or arrange.

  • Edge

    For organizations with stricter governance, control, and on-premises requirements: the same operating loop, with explicit boundaries on the deployment.

What buyers ask

The true answers, including where the answer is “we confirm that during evaluation.”

  • Do we need AI traces?

    No. Forge supports custom model development with or without traces. Existing examples or an agreed effort to build representative examples can be the starting point.

  • What if we do not have a training dataset?

    The first step is to assess the task, the material you have, and whether a suitable example-building effort is feasible. No fixed minimum dataset size is published.

  • Does custom mean training a model from scratch?

    Not necessarily. The primary path is adapting a suitable existing model around an agreed task. Training a foundation model from scratch is not part of Forge’s scope.

  • Do you automatically use our Stack or Edge traffic for training?

    No. Forge uses separately agreed sources and permitted uses. Production payloads on Stack or Edge do not become training material unless you authorize that for a Forge project.

  • Does Forge include ongoing inference hosting?

    No. Forge develops the model. Stack and Edge address inference operations. How a candidate is served is decided with you; it is not included by default.

  • How do we know the result is better?

    By comparing candidates with a baseline against agreed tasks and criteria, including known failures and deployment constraints. Fine-tuning is not itself proof of improvement.

  • Will the model improve continuously?

    Not automatically. Recurring improvement, when offered, needs new approved examples, another evaluation, and an agreed process. A training run does not replace the production model on its own.

  • How long does it take and what does it cost?

    No timeline or price is published. Scope depends on the task, data preparation, method, evaluation, and deployment needs. Book a call and we will walk through it against the job.

See Forge against the job you need done.

Book a Call