Bring your examples
You have approved examples, labels, corrections, or a task dataset. We assess suitability, prepare the data, and build an evaluation baseline.
Create custom and fine-tuned models around your business’s tasks, examples, and requirements. Start with data you already have, or with the task itself. Traces are optional. The result is evaluated against the work it needs to do.
Forge is Sparset’s model-development offering. It takes a defined job, the examples and requirements that describe it, and optional traces, and produces a custom or fine-tuned model evaluated against that job. Training is the method, not the outcome: a candidate ships when it meets the agreed criteria, or the work stops with a record of why it did not. Forge develops the model. Stack or Edge can operate it.
Three starting points that belong in a Forge conversation. Traces are optional. A project still needs representative evidence of the task, and a way to assess results.
You have approved examples, labels, corrections, or a task dataset. We assess suitability, prepare the data, and build an evaluation baseline.
You have relevant interaction records, or you choose to instrument a workflow. Trace collection is optional. We agree the permitted fields and use, then convert approved material into examples.
You know the problem but do not yet have a usable dataset. We assess feasibility and scope a representative example-building effort with you. No traces is a valid starting point; no evidence at all is not a promise.
Forge runs the development process. Your team defines the job, authorizes the data, and accepts or rejects each candidate.
Agree the job, the failures, and what success means. Your team explains the workflow and the constraints.
Turn approved material into training and evaluation examples, kept separate. Your team supplies domain judgment and reviews what may be used.
Choose a supported base model and adaptation method, and train a candidate. Your team resolves task-specific questions; what we can support is confirmed during evaluation.
Compare the candidate with the baseline on agreed tasks, including failures. Your team reviews the evidence and decides whether to accept, iterate, or stop.
Hand over the agreed artifact, the evaluation summary, and the intended-use notes. Your team accepts the deliverable and decides the serving path; a new candidate does not replace production on its own.
See what improved, and what still needs work. The comparison is against the agreed baseline, not a published score.
The candidate is judged against your current approach on agreed tasks and criteria, including cases it was not trained on. A favorable result is an evaluation outcome, not an automatic consequence of fine-tuning.
A model can improve on the target job and get worse in other ways. The report includes representative failures so the acceptance criteria can make those tradeoffs explicit.
Quality is considered alongside the intended latency, throughput, and infrastructure. Compatibility with the serving environment is confirmed for the project, not assumed.
Forge is the model-development offering. Stack and Edge operate inference. You can discuss Forge before choosing a serving environment; bundling is agreed per engagement, not assumed.
For engineering and platform teams that want the model operated in production: planning, deployment, and the operating loop on GPU capacity you own, rent, or arrange.
For organizations with stricter governance, control, and on-premises requirements: the same operating loop, with explicit boundaries on the deployment.
The true answers, including where the answer is “we confirm that during evaluation.”
No. Forge supports custom model development with or without traces. Existing examples or an agreed effort to build representative examples can be the starting point.
The first step is to assess the task, the material you have, and whether a suitable example-building effort is feasible. No fixed minimum dataset size is published.
Not necessarily. The primary path is adapting a suitable existing model around an agreed task. Training a foundation model from scratch is not part of Forge’s scope.
No. Forge uses separately agreed sources and permitted uses. Production payloads on Stack or Edge do not become training material unless you authorize that for a Forge project.
No. Forge develops the model. Stack and Edge address inference operations. How a candidate is served is decided with you; it is not included by default.
By comparing candidates with a baseline against agreed tasks and criteria, including known failures and deployment constraints. Fine-tuning is not itself proof of improvement.
Not automatically. Recurring improvement, when offered, needs new approved examples, another evaluation, and an agreed process. A training run does not replace the production model on its own.
No timeline or price is published. Scope depends on the task, data preparation, method, evaluation, and deployment needs. Book a call and we will walk through it against the job.