Insight

When AI Calls Tools, Where Do Costs Rise?

IT DAILY ·

Jeon Hyeonsang, AI app solutions engineer on Microsoft’s Global Black Belt team

✦ AI Summary

Jeon Hyeon-sang covers where costs rise for AI that calls tools in an IT Daily series.

In a short tool-use task, raising reasoning strength did not keep increasing the pass count, and GPT-5.2 passed 59 out of 60 in 'none' and 60 out of 60 in 'low.'

The article argues that cost cannot be explained by the final answer alone and says call count, results, and termination point must be examined together.

The starting topic of IT Daily's series is where the costs of AI that calls tools begin to rise. The writer is Jeon Hyeon-sang, an AI app solutions engineer on Microsoft’s global Black Belt team.

Jeon Hyeon-sang is working on getting agent, knowledge search, and AI evaluation systems established in real-world operating environments. His career began with signal processing research.

Jeon Hyeon-sang worked on developing a service with 12 million users at SK and also took part in an ML platform build project. At AWS, he spent more than 15 years supporting enterprise AI adoption.

Generative AI is moving from demos to actual business system adoption. However, at adoption sites, concerns are coming up more often than expectations.

In the field, people say cost forecasting is difficult, apart from evaluating reasoning model performance. There are also reactions that even high benchmark scores still have limits in actual in-house tasks.

The gap between performance metrics and the bill, and the gap between demos and practical work, are presented as the biggest obstacles to generative AI adoption. Such issues are also linked to factors that delay adoption even as generative AI enters real work.

This course aims to move the criteria for judging AI use away from intuition and toward measurable metrics, and it also seeks to shift gap detection from guesswork to measurement. To do this, the writer carried out two public experiments. Experiment 1 repeatedly measured the quality, cost, and response time of the same task while adjusting reasoning effort step by step, and examined how operating variables such as cache, retries, and pricing plans affect the bill; in nature, it corresponds to AI bill analysis. Experiment 2 assigned actual work tasks such as document writing and review to AI and then graded the AI output; in nature, it corresponds to AI performance grading. Both experiments converge on the core question of cost per successful result, rather than token usage.

The writer also aimed to provide verifiable evidence instead of slides, and posted all figures together with source data, analysis code, and revision history in a public repository. Measures were taken so that anyone could verify the same method, and failed call records as well as records of missed expectations were also retained. This setup is focused on giving adoption review engineers, architects, and budget decision makers a basis for adoption judgments and budget decisions.

The series topic is evaluating AI against the bill and performance criteria.

The topic for August is when it becomes justifiable to use a reasoning model and comparing models and strengths for multi-step tasks.

The topic for September is the gap between GDPVal RealWorks and benchmark success rates versus actual work success.

The topic for October is the amount of reasoning required when calling tools, agent loop repetition, and hidden costs.

The topic for November is the reliability and cost of GDPVal RealWorks and a file-reading AI grader.

The topic for December is why reasoning remains off by default and the costs and throughput of short fact-based tasks and PAYG and PTU.

The topic for the January 2027 issue is the need for sandbox, multimodal, and long-running environments, and a synthesis of the two experimental groups.

This article is the writer’s personal technical contribution and does not represent the company’s official position. The numerical basis comes from experiment records and aggregated results in the public repository.

In short tool-use tasks, increasing reasoning strength did not keep increasing the pass count.

Earlier issues looked at cost differences and the practical usability of outputs, and this issue focuses on checking the process leading to a result. The August issue examined whether cost differences could arise by model reasoning setting even when the answers were similar, while the September issue checked whether AI completion reports matched results that could actually be used.

What the user sees is one final answer, but the actual process before that final answer is a series of multiple steps. The writer explained that the answer may appear once, but calls can happen multiple times. The issue raised this time was how many times the model is recalled during AI search, calculation, and result reading.

An agent is defined as AI that receives a goal, uses tools, reports the result, and then chooses the next action. This agent can call a calculator after reading search results, or modify a file and then run a command to verify it. In this way, agentic AI moves across tools and goes through multiple steps.

In this process, it is necessary to distinguish the reduction ratio of part of the model input from the reduction ratio of the overall task cost. Even if the answer looks similar, costs can accumulate when model calls accumulate during the process. Costs are also incurred for runs that end before grading.

Therefore, the task result, the reason for termination, and usage need to be recorded together. That makes it possible to trace the intermediate process by which the result is produced and to check what calls and costs accumulated before the final answer was reached.

To understand AI tasks that use tools, model calls and tool execution must be distinguished. A model call means requesting judgment or an answer from the AI service, while tool execution means the actual performance step such as search or calculation. The process consists of the model choosing a tool, receiving the tool result, and recalling the model, and multiple tool execution instructions are possible within a single model response.

When judging cost, the final answer alone is not enough. That is because the model’s resend instructions and previous work results are included in the input. Tokens are the unit used for the model’s text-processing count, and even if the user is shown a short answer, the input and output in the process of reaching the final answer need to be checked separately.

Figure 1 is a conceptual diagram showing four cost-addition points in one tool-using AI task. However, Figure 1 does not observe a specific run, and the effect of compression and cache on cost is only a possible explanation, not a confirmed one.

The experiment presented sample questions. The question content consisted of querying annual revenue for Helio Robotics, querying annual revenue for Vega Logistics, and calculating the combined revenue of the two companies. Korean translation questions were also presented. Helio Robotics and Vega Logistics are fictional companies for the experiment [3].

Rather than the simple addition itself, this experiment treated as the target of verification the procedure of finding the needed information and completing the requested answer through the calculation tool. The items under verification were searching each source, passing data to the calculation tool, and deriving the requested answer. The procedure, including the grading criteria, was to query the two companies separately, add the figures with a calculator, and respond with a single integer [3].

The experiment did not use live internet search. Instead, it used a fixed search tool function configured to return the revenue data for each company. This was intended to reduce the effect of search-result variation by providing the same data.

The experiment also used repeated measurement. Twenty new questions were prepared, and the question topics were whether tools were needed and in what order they should be used. Of these, 6 questions did not require tools, 8 required one tool use, and 6 required multiple tool uses.

The experiment also repeatedly measured whether responses correctly judged the need for tools and the order of use. In addition, the judgment of avoiding unnecessary tool calls was included in the observation target [3].

The experiment was designed to enable comparison by repeating runs under the same conditions. Each question was run 3 times with the same model and settings, GPT-5.2 reasoning strength was adjusted across 5 levels, and a separate comparison baseline was set for GPT-4o. Multiple models were called and evaluated in the same way, and the platform used was Microsoft Foundry.

The total record count was 360. This was calculated as 20 questions × 3 runs × 6 combinations. The number of runs was the same as in the previous article, but the input questions were different from the previous article, and the tool-use procedure was also different. The topic of Table 1 was Measurement Condition 1 - short tool-use experiment.

Figure 2 was a diagram explaining the procedure for processing the revenue query and addition question for the two companies. Figure 2 was organized to separate model judgment from search and calculation tool execution, and it explained the relationship by which the model chooses the next action after receiving the tool result. However, Figure 2 was not a chart measuring the number of calls in a specific run.

Before the experiment, it was expected that as reasoning strength increased in multi-tool connection questions, the number of passes would increase. The basis for this was that a calculation plan would be needed after searching the data. Accordingly, the no-reasoning setting was seen as unfavorable, and it was expected that once a high level was reached, the additional quality gain would be small but the cost would rise.

Looking first at the overall results, GPT-5.2 passed 59 out of 60 in 'none' and 60 out of 60 in 'low.' The no-reasoning setting also showed a high pass count. Raising the strength level step by step did not lead to a step-by-step improvement in results. The overall results for the 6 combinations, including GPT-4o, are shown in Figure 3.

The initial hypothesis concerned multiple-tool connection questions. The number of questions selected in advance was 6, and each question was run 3 times. Accordingly, the analysis target for that question group was 18 runs.

In this question group as well, GPT-5.2 passed 17 in 'none' and 18 in 'low.' The gap between 'none' and 'low' in that question group was 1. Even in this question group, further increases in strength did not increase the number of passes.

In the advance notice wording, it is correct to say that the first value-recording subject for the experiment was 'low' rather than 'none,' but that wording should not be generalized to the overall results as they are, so the scope of interpretation needs to be narrowed. In this sample, 'low' was the lowest reasoning strength that passed all responses, and 'none' passed almost all responses. However, a 1-run difference alone cannot justify the conclusion that low reasoning strength is always more accurate or more economical for all tool tasks.

Cost was defined not as the actual billed amount but as a calculated value obtained by multiplying the input and output usage reported by the model service by a fixed price table. Here, 'API calculation cost' means that calculated value, and API is the interface through which a program calls the model service. Therefore, the calculated cost is separate from the amount reconciled against the actual bill.

The total cost of each setting was divided by the 60 runs performed, regardless of pass or fail, and this average cost was converted to a 1,000-run task basis. This 1,000-run basis is not the actual value for 1,000 runs, but a converted value for comparison in the same unit. On that basis, GPT-5.2 'none' was about USD 2.76, GPT-5.2 'low' was about USD 2.84, and GPT-5.2 'very high' was about USD 3.50. In this grading, 'low' and 'very high' both passed, and there was a cost difference between the two settings.

The comparisons of cost and pass count were both presented based on the sample internal average and judgment values. The cost comparison basis was the average per run, and the pass basis was the GPT-4o grader’s 2-point judgment. In the results from this sample, 'none' had the lower cost, and in the results from this sample, 'low' had 1 more passing response. Figure 3 contains the pass count and API calculation cost for the short tool-use experiment, and the cost formula was the average converted to 1,000 task runs across all questions. The cost tally also included one exceptional run for 'low'.

The recounted basis was 360 runs in total, 60 runs for each of the 6 combinations. The multi-tool question group consisted of 6 questions classified in advance, and the multi-tool question group was run 3 times per question. The number of multi-tool question-group runs per combination was 18. In this sample, 'none' was cheaper and 'low' gained 1 more pass, but the basis for using the results was to check the business necessity of the extra cost rather than to decide a winner.

Also, the entity that judged the pass count was the GPT-4o grader. In the public materials, there was no score for agreement with human grading. Accordingly, the scope of interpretation of the results was limited to that question and that grading standard. The limitation of the pass-count meaning was also that it did not indicate the success probability outside the sample.

The earlier experiment had the characteristics of short length and a limit on the number of tool interactions. Accordingly, cases such as file modification and command execution, where the process becomes long-running, were presented, and a further question was raised as to whether the resubmitted content could be reduced for the next judgment and whether the task could still be completed normally after reducing it.

This question was examined through a separate input-compression experiment. Here, prompt means the instructions and materials delivered to the model, and prompt compression means shortening part of the prompt. The range of compression candidates was set to file lists and installation records among the outputs returned by the tool, and candidates were selected based on sections that passed preset rules. By contrast, task instructions, code, structured data, code-description mixed sections, and sections with unclear boundaries were left uncompressed.

However, a caution was also given that output classified as a candidate does not guarantee the safety of discarding internal information [7][14].

The task was taken from the public evaluation set Terminal-Bench 2.1, which is a set for command-based computer task execution and result checking. The comparison was divided into the method solved without additional compression, the squeez-applied method, the Headroom-applied method, and the LLMLingua-2-applied method. LLMLingua-2 is a compression model released by Microsoft and Tsinghua University researchers. Each method started in a fresh workspace for the same task, and Headroom used a limited setting centered on repeated file-path grouping.

Next, to determine what was reduced based on the file list, a separate inspection was presented in which a compressor was applied to saved tool output in inspection data separate from the 104 runs. This public example was based on an actual inspection case, and private paths were replaced with <LOG_DIR> for publication. In the separate inspection result, Headroom preserved the same folder path only once.

In the example, it was checked whether the original location could be restored even if the common path was removed and only the file name was left.

The 3 tool-output items before compression were <LOG_DIR>/2025-07-03_api.log, <LOG_DIR>/2025-07-03_app.log, and <LOG_DIR>/2025-07-03_auth.log.

The 3 strings after compression were 2025-07-03_api.log, 2025-07-03_app.log, and 2025-07-03_auth.log.

The target of this reduction was the repeated path, and the information retained was each file name and the information needed to restore the location.

A separate inspection was also performed.

In this inspection, it was checked whether the restored string matched the original, and byte-level verification was also performed.

However, in other compression-setting cases, there was loss in the latter part of the installation output, and there was also loss of words signifying error or refusal.

For that reason, input shortening alone could not be used to determine whether the necessary information had been preserved.

The scope of this inspection was checking changes in the string, and the unmeasured item was whether that input made the model perform the task better [14].

The author stated that 104 records were collected by running 1 run each across 4 methods for 26 tasks, with all 4 conditions executed and verified. The author then said that 40 of the 104 records passed the task-embedded check, and explained that the embedded check was the judgment of the verification program included in the task. The author also said that the embedded check and the model’s self-score are separate, and that the same judgment was confirmed by restoring the workspace and rechecking it.

However, the author said that the nature of this data was not to rank compression tools. Instead, the data was described as a preliminary observation to identify points needing re-verification. The task-selection criteria were the presence of compressible output in past records and the order of a pre-established task list, and it was also added that the sample was not a random sample representing all work. In addition, it was explained that the number of runs per task and condition was 1.

Also, in the earlier separate inspection, the same task was run repeatedly without compression, and the pass judgment changed even in those repeated runs without compression. The range of results in the first half and second half of the repeated measurements did not satisfy the pre-set stability condition, and accordingly the comparison criteria before and after compression were themselves unstable. The author said that repeated measurement was needed, along with additional controls, if one wanted to judge only the effect of compression from the differences shown below.

The topic of Table 2 is the preliminary experiment on input compression under Measurement Condition 2. Here, it is important to distinguish between reduced input and reduced cost as two different numbers. If input is shortened, one may assume cost also falls by the same ratio, but that needs verification. Therefore, the items to check are the location of the reduction and the amount of the reduction.

Accordingly, the parts that actually changed were first tallied separately. There were 78 applications of additional compression, and 23 actual string changes. The change spans whose before-and-after originals could be confirmed numbered 209. The token-count basis used was the same token counter mentioned earlier. The number of tokens in these change spans fell by about 48%. The subject of the about 48% reduction rate is the change spans.

By contrast, the model input and cost of the overall task include factors beyond the change spans, so separate calculation is needed. The elements included in the separate calculation are instructions, history, and output. Therefore, the reduction rate confirmed in the change spans and the model input and cost of the overall task cannot be treated as the same item.

The actual recorded results are presented for four conditions. The cost tally below is the total of the 26 runs in each condition that reached grading. This tally includes the cost of runs that did not pass.

This observation has the character of a preliminary observation made once per task and condition. Accordingly, it is not a ranking table for compression tools.

The subject under review is the unit price of input calculation. The prompt cache function reuses the same prefix of previously processed input. Looking at the sum of input tokens for 26 runs per condition, the cache-input share was about 64.5% for no additional compression and about 46.7% for Headroom. However, even though the cache-input share declined, the API calculation cost also declined from USD 5.62 to USD 5.36. Therefore, this case cannot be interpreted as a case where cost rose because the cache-input share fell.

The basis for reviewing cache together lies in the pricing structure. Under the fixed price table used in this experiment, general input cost USD 2.50 per 1 million tokens, and cache input cost USD 0.25. The unit price of cache input is one-tenth that of general input. Because of this difference, the share of cache input among input tokens and the actual API cost must be compared together, and cost cannot be determined from the change in share alone.

One possible explanation is that when input is reduced through compression, the reused prefix changes and the cache discount may shrink. In that case, if the reduction in the cache discount occurs as well, the cost may not fall by as much as the token reduction. If the lost discount is larger, cost could even rise. However, this is only a possible explanation of cost change, not a confirmation of the cause of the Headroom result.

In this observation, it was impossible to isolate effects because of differences in the number of requests per condition, differences in the model task path, uncontrolled cache effects, and uncontrolled concurrent execution effects. The observation showed cost differences, but Headroom’s lower cost alone could not prove the cost-saving effect of compression. Accordingly, to answer the question of whether reducing input also reduces task cost, it was necessary to check quality, usage, and cost at the same time for each condition, and to confirm whether the differences persisted in repeated runs.

Across all 104 completed grading runs, model API call attempts totaled 1,027, and the API calculation cost was about USD 22.33. This figure was the sum of the costs for 40 passing cases and 64 failing cases. The meaning of that total was to show the scale of recorded runs.

Costs were also incurred for runs that ended before grading. In addition to the table of completed grading results, there were also runs that were separately reviewed. Therefore, the flow required looking not only at the results in the table, but also at the runs outside it.

In addition, selecting the experimental tasks took longer than expected. The reason was that there were too few items meeting the screening criteria. In that process, the limits on calls, time, and cost were lifted so that screening and evaluation could proceed.

Five runs recorded separately were 5 separate recorded attempts, and each failed to reach the first grading even after about 10 hours. These 5 runs were recorded according to the observation criteria, and they were not originally designed to run for a fixed 10 hours. After checking the execution state, the operator ended them, and the 5 runs came from 4 different tasks. All 5 ended before grading, so quality remained unconfirmed, and whether they would have been resolved with additional waiting was also unknown. Whether the output passed inspection was also unknown.

Looking at the termination categories, 4 cases were operator-initiated terminations and 1 case was a stall while waiting for a response. In this way, the reasons for termination were divided between interruption by operator judgment and a stall while waiting for a response.

The 5 runs had 2,783 model API call attempts, and the API calculation cost based on responses with confirmed usage was about USD 87.77. The estimated cost for 1 request with unconfirmed usage was excluded from the total, and the cost calculation basis was the same unit price and formula as the earlier USD 22.33. However, because the USD 87.77 group was a separate group with different tasks and execution conditions, performance comparison is not possible between USD 22.33 and USD 87.77, and compression-effect comparison is also not possible.

Usage was confirmed, but quality was still unconfirmed, so cost and quality remain in separate states. If this state is treated as an error, the quality statistics change. Conversely, if only completed results are collected and records are excluded, the cost disappears. Therefore, cost must be preserved, and cost and grading results need to be distinguished.

Figure 4 presents the 104 completed grading records and the 5 separately attempted pre-grading termination records. However, the two sections include different tasks and different execution conditions. Accordingly, performance comparison between the two sections was not carried out, and compression-effect comparison between the two sections was also not carried out. Also, API calculation cost does not include compressor computer costs or verification program computer costs, and it is not the reconciled amount on the actual bill. One request with unconfirmed usage among the 5 long-running attempts was excluded from the cost total. The references are [8][10][12].

After the records, the conditions for starting the next paid run were changed. The model API call limit for one attempt was set to a maximum of 60, including retries. The execution-time limit, including preparation, work, and grading, was set to a maximum of 40 minutes. A cost ceiling per attempt was pre-set and approved, and the overall execution cost ceiling was also pre-set and approved.

The termination time was also pre-set and approved. If a required value was blank, the run was stopped before model calls were made. Such criteria are shown under reference [13].

When execution reaches a limit, the records must be separated. At that point, whether the call-count and cost limits were reached, and whether the time and other execution limits were reached, must be recorded as the reason for stopping. If it stops before grading, the quality state must be treated as unconfirmed. In addition, the last call position, confirmed usage, confirmed cost, and unconfirmed cost must be preserved.

The policy values for this execution run are 60 calls and 40 minutes. However, determining whether the policy values are appropriate for each task needs to be done separately. The current verification scope is the new policy and implementation. Whether repeated use of the new standard for the same task changes cost or hinders obtaining the needed result has not yet been verified.

The operator should review the definition of completion criteria, the permitted time and call counts while waiting for results, what must be preserved when limits are reached, and who is responsible for follow-up checks after limits are reached. In addition, execution conditions need to be determined together with reasoning-strength selection and compression-tool selection.

The three earlier cases show that when looking at the cost of tool-using AI, it is difficult to explain it with only the final answer or a single setting. In the short tool-use experiment, reasoning strength and results were checked at the same time. In the input-compression preliminary experiment, the reduced section size and the overall usage showed different numbers. In the pre-grading termination runs, it was confirmed that costs need to be recorded even when there is no result.

For that reason, when judging cost, the number of calls after the task starts, the results left behind, and the termination point should be examined together. It is also necessary to distinguish completed results from pre-grading terminations. Once this distinction is applied, the additional measurement items needed when separating completed results from pre-grading terminations can also be made more specific. The next item to verify is the judgment itself.

In this article, passing means satisfying the GPT-4o grader or the task-embedded verification standard. When reviewing the meaning of the judgment standard next, it is necessary to distinguish between whether the same test result can be reproduced and whether that test accurately separates actual work quality. Those are not the same problem.

In the separate baseline of the compression experiment, there were also cases where the embedded check rejected equivalent expressions. This leads to the point that additional verification is needed when interpreting embedded check results. The related references are [9][15].

The next issue will cover an AI grader that reads files and evaluates outputs.

This article narrows the questions for the next installment to the reliability of the reasons an output passed and the cost of confirming the judgment.

Reference numbers are used to connect the basis of the draft.

The fixed standard for repository materials is the commit that confirms the writing time.

[1] refers to Jeon Hyeon-sang, IT Daily, 'When Will Reasoning Models Be Worth Their Price?', 2026-07-29.

The source of the preview sentence in the body of [1] is the wording in the author-confirmed published version, and [2] refers to IT Daily, the published version of the second installment in the series.

[3] refers to the tool-use experiment questions, structure, and original question text in When Reasoning Pays Off, and [4] refers to the experiment setup, execution report, and pricing table at the time in When Reasoning Pays Off.

[5] refers to the per-run analysis data, raw usage records, and grading records in When Reasoning Pays Off, and the method used in the body of [5] to calculate the cost for 360 runs was a recalculation that applied the [4] pricing table to usage per run before exclusions.

The listed items summarize which parts were used from each reference and the scope of application for calculation and citation.

[6] uses the auxiliary statistics from When Reasoning Pays Off and the scope where human grading is absent.

[7] uses the outline and measurement conditions of the preliminary experiment in Prompt Compression Billing Bench, and [8] uses the aggregated JSON of the preliminary experiment in Prompt Compression Billing Bench.

In particular, the cache share is calculated by dividing the cache input tokens in provider_usage.by_condition in [8] by the total input tokens.

[9] uses the review of the experiment design and the limitations of the judge in Prompt Compression Billing Bench, and [10] uses the cost formula and aggregation scope of Prompt Compression Billing Bench.

The explanation related to prompt caching is sourced from Microsoft Learn’s 'Prompt caching' in [11], with the verification date being 2026-09-21, and only the explanation of the matching prefix of the input and prompt-cache reuse is used in a limited way. The fixed unit-price basis for the experiment also follows [10].

[12] through [15] are the inspection and record items that continue in Prompt Compression Billing Bench from post-operation handling after the operator ended the run, paid-run policy, stored-output inspection, and repeated-run stability. [12] includes long-tail content from the technical basis section of Prompt Compression Billing Bench that was terminated later by the operator, and also includes the release of restrictions and the follow-up decision record of Prompt Compression Billing Bench.

[13] records the termination policy and cost policy for the next paid run in Prompt Compression Billing Bench, and also includes the policy-check implementation in Prompt Compression Billing Bench.

[14] contains the separate inspection in which a compressor was applied to saved output for Prompt Compression Billing Bench, the pre- and post-release example presentation in Prompt Compression Billing Bench, and the distinction between compression candidates and protected targets in Prompt Compression Billing Bench. [14] also distinguishes the scope of release and the nature of the data by stating that the examples and paths are characteristics of the example, based on static inspection, and that the 104-run results are not written as original text.

[14] also states that the path examples are a processed public version of the source data and that the source of the path examples in [14] is a static inspection without model calls. In addition, the 104 task runs in [14] were not written in the original text.

[16] is an external basis for LLMLingua-2, presenting the LLMLingua-2 section of Microsoft Research’s Research Focus: Week of April 15, 2024. [16] additionally states that LLMLingua-2 is a joint study by Microsoft and Tsinghua University researchers, that the reference for the model and code release is Microsoft’s LLMLingua repository, and that the verification date is 2026-09-21.

The execution date for the tool-use experiment was 2026-05-24, and the execution window for the input-compression task comparison was 2026-09-16 to 17 (UTC). The deployment for model calls was Azure on Microsoft Foundry, the model used was GPT-5.4, the provider-reported version was gpt-5.4-2026-03-05, the generation settings were temperature=0 and reasoning_effort=none, the compression tools used were the file-path-only settings of squeez 1.48.4 and Headroom 0.36.5, and LLMLingua-2 0.2.2, the basis for change-span calculation was tiktoken 0.14.0’s o200k_base, and those settings alone do not guarantee determinism or representativeness of the outputs.

Source: IT DAILY · Jeon Hyeon-sang
Original: https://www.itdaily.kr/news/articleView.html?idxno=241904

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.