top of page

AI Usage Skill Visualizer
Now Available in Beta

Extract AI logs in three clicks.

About the Overall Evaluation

The overall evaluation considers not only the results of the five evaluation dimensions, but also the processes of decision-making, selection, and revision shown in the AI conversation log, and how those processes are reflected in the final output.

Rather than viewing each evaluation dimension independently, it examines the AI interaction as a whole to understand how decisions and revisions made toward the intended goal are connected and ultimately reflected in the final output.

1. Logical Quality of the Output

What is evaluated:
The structure and logical development of the final output, the connection between supporting reasons and conclusions, and whether any contradictions are present.

Evaluation approach:
The evaluation looks beyond the surface quality of the output and examines whether information is appropriately organized for the intended purpose and whether the output is built on consistent reasoning.

What a strong result looks like:
The purpose and conclusion are clear, the reasoning, structure, and content are consistent, and unnecessary logical leaps or contradictions are minimal.

2. Decision-Making Process

What is evaluated:
Whether the user makes decisions about AI suggestions through comparison, selection, acceptance, or rejection.

Evaluation approach:
The evaluation examines whether the user considers alternatives in relation to the intended purpose and conditions, rather than simply accepting AI responses, and determines the direction through their own judgment.

What a strong result looks like:
The user compares multiple options, explains the reasons for accepting or rejecting them, and continues to make decisions aligned with the intended purpose.

3. Revision Process

What is evaluated:

Whether the user identifies problems in AI responses or intermediate results, makes specific revisions, and improves the final output.

Evaluation approach:

The evaluation looks beyond simple regeneration and examines whether the user identifies what is wrong and improves the result by adding conditions, changing direction, restructuring, or making other targeted revisions.

What a strong result looks like:

The user clearly identifies specific problems, has a clear intention behind each revision, and those revisions lead to meaningful improvements in the final output.

4. Human Agency

What is evaluated:

Whether the user, rather than the AI, takes the lead in defining the purpose, conditions, evaluation criteria, and final decisions.

Evaluation approach:

The evaluation is not based on how many instructions the user gives to the AI. It examines whether the user defines what they are trying to achieve, manages the AI’s suggestions, and moves the work forward by rejecting or modifying those suggestions when necessary.

What a strong result looks like:

The purpose and decision criteria remain with the user, who selects and revises AI suggestions while making the final decisions about the direction of the work.

5. Overall Coherence

What is evaluated:

The connection between the decisions, selections, and revisions shown in the conversation log and the final output.

Evaluation approach:

The evaluation examines whether the intentions and judgments expressed during the conversation are actually reflected in the final output, and whether the overall AI usage process remains consistent without contradictions.

What a strong result looks like:

The decisions and revisions made during the conversation are clearly reflected in the final output, and there is a consistent flow from purpose → conversation → judgment → revision → output.

Scroll

bottom of page