Loading...

Sber unveils GigaChat 3.5 Reasoning for complex enterprise workflows

Sber unveils GigaChat 3.5 Reasoning for complex enterprise workflows

The Russian banking and technology group’s latest AI model combines specialised reasoning training with an extended context window, targeting applications in software development, analytics and business operations.

Sber, the Russian banking and technology leader, has introduced GigaChat 3.5 Reasoning, expanding its artificial intelligence offering with a model designed to plan, verify and work through complex tasks. Announced on September 10, 2026, the release is aimed at scenarios where a useful answer depends on several connected decisions, from resolving software errors to analysing documents and preparing  resource-constrained plans.

According to the company, the model can break a problem into stages, identify missing information, use search and other tools when needed, and check intermediate results before responding. It can also revisit an earlier step if its proposed solution fails to meet the task’s requirements. Sber reports the largest gains over GigaChat 3.5 Ultra Instruct in mathematics, coding, instruction following and structured output.

For enterprises, these capabilities have a practical relevance. A software fix must work within the surrounding code, a schedule must respect budgets and deadlines, and a customer-service response may require information from several systems. GigaChat’s new reasoning mode is designed to support the steps between receiving such a request and producing a usable result.

Loading...

Bringing reasoning into business workflows

Software development is a central use case. The model’s training covers algorithms, code editing and test generation, alongside repository-level tasks in which a proposed patch is applied and tested. This gives development teams a basis for evaluating assistance with debugging and maintenance, where understanding dependencies and checking the effects of a change are essential parts of the job.

Document analysis offers another potential application. Sber identifies the ability to look for inconsistencies between contract clauses as an example of a task requiring connected reasoning. The model could help surface passages for further review by comparing requirements across a document and identifying where they appear to conflict. 

In logistics and planning, the company points to routes and schedules that must account for deadlines, budgets and available resources. A travel-planning example illustrates the sequence: clarify missing preferences, select flights and accommodation, check the combined budget and adjust the plan if the options do not fit. The same pattern of gathering information, testing constraints and revising choices is relevant to operational planning.

Customer support and analytics are also among the proposed applications. A support assistant could clarify a request and consult different systems before responding, while an analytical workflow could combine calculations with successive checks of a hypothesis. These are  intended use cases; their value in deployment will depend on the data, tools and controls available to the system.

Anton Frolov, senior vice president and head of GenAI Development at Sberbank, said the ability to identify missing data, validate intermediate conclusions and modify a solution was particularly important in programming, analytics and agent-based scenarios, where outcomes depend on a sequence of decisions.

Six specialised experts behind one model

GigaChat 3.5 Reasoning is the first GigaChat model with full reasoning trained using online reinforcement learning, according to its technical documentation. Post-training begins from a supervised fine-tuning checkpoint. Six domain experts are then trained separately before their capabilities are combined into the release model.

The experts cover science and mathematics, coding, code agents, general agents, dialogue, and capabilities such as instruction following, long-context handling and structured output. Training them independently allows the developer to use different feedback mechanisms for different types of work.

Mathematical solutions are assessed through final-answer verification, while coding tasks are evaluated by executing code. Repository-level changes are checked through tests after a patch is applied. General agent tasks are evaluated against the final state of the environment, while dialogue responses undergo side-by-side assessment by an AI judge.

The training workload becomes more demanding as the model improves. Before training, the current checkpoint is evaluated against the task pool, and problems solved in more than 75% of attempts are removed. Rewards combine checks for mandatory constraints with measures of answer quality and an adaptive length penalty.

The six experts are brought together through on-policy distillation. In this process, the student model generates its own solution trajectory, and the relevant expert provides token-level supervision along that path. The approach is intended to transfer specialised capabilities into a single model that can handle a broader range of requests.

Architecture for larger inputs

GigaChat 3.5 Reasoning uses a Mixture-of-Experts architecture with 432 billion total parameters and 28 billion active parameters. It combines Multi-head Latent Attention with GatedDeltaNet linear-attention layers in a custom hybrid design. Sber says the linear-attention component helps the model process long contexts more efficiently.

The model supports a maximum context length of 262K tokens, creating room for substantial documents, extended conversations or collections of code within an interaction. For business applications, this capacity is relevant when a task requires information spread across many passages or files. Context capacity describes how much input the model can accommodate; it does not, by itself, establish the accuracy of its analysis.

The model was trained natively in FP8 precision at all stages, with FP8 weights provided for inference. A separate BF16 version is available for fine-tuning or custom quantisation. The architecture also includes three multi-token prediction heads for speculative decoding, a feature intended to support efficient generation.

Loading...

Reported gains in planning and coding

Sber reports that the model’s IFBench score increased from 44 to 77 compared with the non-reasoning version. Its Natural Plan score rose from 64 to 80, while Live Code Bench v6 increased from 56 to 85. These evaluations cover instruction following, planning and programming, respectively, making them relevant to the workflows the company is targeting. 

The company also reports that GigaChat 3.5 Reasoning uses an average of 37% fewer tokens on mathematical problems than DeepSeek V4 Flash Preview. This is a developer-reported comparison for a particular workload; it should not be read as an equivalent reduction in overall operating costs or response times.

Users can activate reasoning mode when needed, according to Sber. Simple questions can receive immediate answers, while complex problems may require tens of thousands of tokens to work through. For organisations evaluating the model, the relevant balance will be
between task accuracy, response time and the resources required to produce a result.

Developer access and the enterprise opportunity

Sber says GigaChat 3.5 Reasoning is available to developers, with model weights released on Hugging Face under the MIT licence. At launch, the company also said business access through an API would follow. Access to the weights gives developers scope to assess the model against their own workloads and explore integration into their products.

With specialised reasoning training, an extended context window and developer access, Sber is positioning GigaChat 3.5 Reasoning for applications spanning software development, document analysis and customer operations. The next step for enterprise adoption is translating those capabilities into workflows that can complete tasks consistently within the organisation’s requirements for accuracy, speed and cost.


Sign up for Newsletter

Select your Newsletter frequency