ij commited on
Commit
f3d5a27
·
verified ·
1 Parent(s): 072ed2a

Publish calibrated LoRA detector and model card

Browse files
README.md ADDED
@@ -0,0 +1,153 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ko
4
+ license: gemma
5
+ library_name: peft
6
+ pipeline_tag: text-classification
7
+ base_model: google/embeddinggemma-300m
8
+ base_model_relation: adapter
9
+ tags:
10
+ - text-classification
11
+ - ai-text-detection
12
+ - authorship-analysis
13
+ - korean
14
+ - fiction
15
+ - stylometry
16
+ ---
17
+
18
+ # ⚠️ Do not use this model as evidence of AI authorship.
19
+
20
+ **Munche-768-AI-Detector was built for dataset triage and proof-of-concept research. It cannot establish plagiarism, copyright infringement, misconduct, or whether a person or an AI wrote a text. Its output can be wrong.**
21
+
22
+ # Munche-768-AI-Detector
23
+
24
+ <p align="center">
25
+ <img src="./assets/baragi-ai.png" width="128" alt="Baragi AI">
26
+ </p>
27
+
28
+ Munche-768-AI-Detector classifies Korean genre-fiction passages as `human`, `uncertain`, or `llm`. It starts from [Munche-768](https://huggingface.co/Baragi-AI/Munche-768), then jointly tunes LoRA weights in the top four Transformer layers and a 769-parameter linear classifier.
29
+
30
+ ## Overall test result
31
+
32
+ The sealed test combines 277 human passages and 256 LLM passages from the independent-generation and content-preserving rewrite evaluations.
33
+
34
+ | Metric | Result |
35
+ |---|---:|
36
+ | AUROC | **98.59%** |
37
+ | Binary accuracy | **94.00%** |
38
+ | Balanced accuracy | **93.91%** |
39
+ | Human recall | **96.03%** |
40
+ | LLM recall | **91.80%** |
41
+ | Human false-positive rate | **3.97%** |
42
+
43
+ <p align="center">
44
+ <img src="./assets/overall-confusion-matrix.svg" width="900" alt="Overall binary confusion matrix across independent generation and content-preserving rewrites">
45
+ </p>
46
+
47
+ ## Independent-generation test
48
+
49
+ The sealed test contains 87 passages from human-written novels and 66 passages written directly by 11 language-model families. Content-preserving rewrites are excluded from this evaluation.
50
+
51
+ | Metric | Result |
52
+ |---|---:|
53
+ | Binary accuracy | **100.00%** |
54
+ | Balanced accuracy | **100.00%** |
55
+ | Human recall | **100.00%** |
56
+ | LLM recall | **100.00%** |
57
+ | Human false-positive rate | **0.00%** |
58
+
59
+ <p align="center">
60
+ <img src="./assets/independent-confusion-matrix.svg" width="900" alt="Binary confusion matrix for independently written human and LLM fiction">
61
+ </p>
62
+
63
+ ## Three-way decision
64
+
65
+ The thresholds were selected on validation data only. The classification score is not a calibrated probability that a passage was written by AI. Results below cover the full sealed test.
66
+
67
+ | Score | Output |
68
+ |---:|---|
69
+ | `≤ 0.5000` | Human |
70
+ | `0.5000 < score < 0.8697` | Uncertain |
71
+ | `≥ 0.8697` | LLM |
72
+
73
+ | Metric | Result |
74
+ |---|---:|
75
+ | Coverage | **93.62%** |
76
+ | Accuracy among classified passages | **95.19%** |
77
+ | Human classified as LLM | **1.08%** |
78
+ | Human classified as uncertain | **2.89%** |
79
+ | LLM classified as human | **8.20%** |
80
+ | LLM classified as uncertain | **10.16%** |
81
+
82
+ Within the independent-generation subset, all 87 human passages received a human decision. Of the 66 LLM passages, 64 received an LLM decision and two were uncertain.
83
+
84
+ <p align="center">
85
+ <img src="./assets/three-way-outcomes.svg" width="900" alt="Human, uncertain, and LLM outcomes for the independent-generation test">
86
+ </p>
87
+
88
+ ## Content-preserving rewrite test
89
+
90
+ This test contains 190 human passages and 190 LLM rewrites that preserve the source content. It is harder than distinguishing independently written human and LLM fiction.
91
+
92
+ | Metric | Binary decision | Three-way decision |
93
+ |---|---:|---:|
94
+ | AUROC | **97.47%** | N/A |
95
+ | Balanced accuracy | **91.58%** | N/A |
96
+ | Human recall | **94.21%** | N/A |
97
+ | LLM recall | **88.95%** | N/A |
98
+ | Human classified as LLM | 5.79% | **1.58%** |
99
+ | LLM classified as human | 11.05% | **11.05%** |
100
+ | Coverage | N/A | **91.58%** |
101
+ | Accuracy among classified passages | N/A | **93.10%** |
102
+
103
+ <p align="center">
104
+ <img src="./assets/benchmark-comparison.svg" width="900" alt="Performance comparison between independent generation and content-preserving rewrites">
105
+ </p>
106
+
107
+ ## Input length
108
+
109
+ Use passages between **384 and 2,048 EmbeddingGemma tokens**. Inputs shorter than 384 tokens are not supported as stable operating inputs. Inputs longer than 2,048 tokens must be divided into separate windows before classification.
110
+
111
+ The training and evaluation corpora covered short rewrite passages near 400 tokens and independent fiction passages near the 2,048-token model limit. Document-level aggregation across multiple windows has not been calibrated.
112
+
113
+ ## Usage
114
+
115
+ Access to the gated EmbeddingGemma base model is required.
116
+
117
+ ```bash
118
+ pip install torch numpy sentence-transformers peft safetensors
119
+ ```
120
+
121
+ ```python
122
+ from inference import MuncheAIDetector
123
+
124
+ detector = MuncheAIDetector(".")
125
+ result = detector.predict(korean_fiction_passage)
126
+
127
+ print(result)
128
+ # {"label": "human" | "uncertain" | "llm", "score": float, "tokens": int}
129
+ ```
130
+
131
+ ## Training data
132
+
133
+ | Split | Human | LLM |
134
+ |---|---:|---:|
135
+ | Train | 1,315 | 1,135 |
136
+ | Validation | 290 | 275 |
137
+ | Test | 277 | 256 |
138
+
139
+ The training set combines human-written Korean genre fiction, independently generated LLM fiction, and content-preserving LLM rewrites. GPT-5.6 Sol Medium contributes 24 training passages and six validation passages. No Sol Medium passage was added to test.
140
+
141
+ The detector was initialized from Munche-768. Only LoRA weights in Transformer layers 20-23 and the linear classifier were updated. The selected checkpoint is step 275. A preservation loss limited movement away from the original Munche-768 embedding during tuning.
142
+
143
+ Raw human fiction is not distributed with this repository.
144
+
145
+ ## Limitations
146
+
147
+ - The model was trained and evaluated on Korean genre fiction.
148
+ - Generalization to language-model families absent from training remains unknown.
149
+ - The model cannot identify text jointly written or substantially edited by humans and AI.
150
+
151
+ ## License
152
+
153
+ Munche-768-AI-Detector is derived from `google/embeddinggemma-300m` and Munche-768. Use is subject to the Gemma license and the access terms of the gated base model.
adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "google/embeddinggemma-300m",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": false,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 64,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.05,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 32,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "gate_proj",
34
+ "down_proj",
35
+ "v_proj",
36
+ "up_proj",
37
+ "o_proj",
38
+ "k_proj",
39
+ "q_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "FEATURE_EXTRACTION",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8802f80c3dc24b370dc6cb2937ac559475220ca4d32c7ce34a5664432cc4598f
3
+ size 33465560
assets/baragi-ai.png ADDED
assets/benchmark-comparison.svg ADDED
assets/independent-confusion-matrix.svg ADDED
assets/overall-confusion-matrix.svg ADDED
assets/three-way-outcomes.svg ADDED
calibration.json ADDED
@@ -0,0 +1,78 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "selection": {
3
+ "split": "validation",
4
+ "human_as_ai_max": 0.05,
5
+ "human_boundary": "fixed at the binary threshold"
6
+ },
7
+ "thresholds": {
8
+ "human_max": 0.5,
9
+ "ai_min": 0.8696600198745728
10
+ },
11
+ "validation": {
12
+ "combined": {
13
+ "samples": 565,
14
+ "coverage": 0.9079646017699115,
15
+ "covered_accuracy": 0.9551656920077972,
16
+ "human_as_ai": 0.04827586206896552,
17
+ "human_uncertain": 0.08275862068965517,
18
+ "ai_as_human": 0.03272727272727273,
19
+ "ai_uncertain": 0.10181818181818182
20
+ },
21
+ "rewrite": {
22
+ "samples": 406,
23
+ "coverage": 0.8768472906403941,
24
+ "covered_accuracy": 0.9353932584269663,
25
+ "human_as_ai": 0.06896551724137931,
26
+ "human_uncertain": 0.11330049261083744,
27
+ "ai_as_human": 0.04433497536945813,
28
+ "ai_uncertain": 0.1330049261083744
29
+ },
30
+ "independent": {
31
+ "samples": 153,
32
+ "coverage": 0.9934640522875817,
33
+ "covered_accuracy": 1.0,
34
+ "human_as_ai": 0.0,
35
+ "human_uncertain": 0.011494252873563218,
36
+ "ai_as_human": 0.0,
37
+ "ai_uncertain": 0.0
38
+ },
39
+ "sol_medium": {
40
+ "samples": 6,
41
+ "coverage": 0.8333333333333334,
42
+ "covered_accuracy": 1.0,
43
+ "human_as_ai": null,
44
+ "human_uncertain": null,
45
+ "ai_as_human": 0.0,
46
+ "ai_uncertain": 0.16666666666666666
47
+ }
48
+ },
49
+ "test": {
50
+ "combined": {
51
+ "samples": 533,
52
+ "coverage": 0.9362101313320825,
53
+ "covered_accuracy": 0.9519038076152304,
54
+ "human_as_ai": 0.010830324909747292,
55
+ "human_uncertain": 0.02888086642599278,
56
+ "ai_as_human": 0.08203125,
57
+ "ai_uncertain": 0.1015625
58
+ },
59
+ "rewrite": {
60
+ "samples": 380,
61
+ "coverage": 0.9157894736842105,
62
+ "covered_accuracy": 0.9310344827586207,
63
+ "human_as_ai": 0.015789473684210527,
64
+ "human_uncertain": 0.042105263157894736,
65
+ "ai_as_human": 0.11052631578947368,
66
+ "ai_uncertain": 0.12631578947368421
67
+ },
68
+ "independent": {
69
+ "samples": 153,
70
+ "coverage": 0.9869281045751634,
71
+ "covered_accuracy": 1.0,
72
+ "human_as_ai": 0.0,
73
+ "human_uncertain": 0.0,
74
+ "ai_as_human": 0.0,
75
+ "ai_uncertain": 0.030303030303030304
76
+ }
77
+ }
78
+ }
config.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architecture": "EmbeddingGemma with Munche-768 initialization, detector-tuned LoRA, and linear head",
3
+ "base_encoder": "google/embeddinggemma-300m",
4
+ "initial_adapter": "Baragi-AI/Munche-768",
5
+ "embedding_dim": 768,
6
+ "labels": {
7
+ "0": "human",
8
+ "1": "llm"
9
+ },
10
+ "max_tokens": 2048,
11
+ "supported_token_range": [384, 2048],
12
+ "selected_step": 275,
13
+ "updated_lora_layers": [20, 21, 22, 23],
14
+ "trainable_lora_parameters": 1392640,
15
+ "trainable_head_parameters": 769,
16
+ "seed": 20260727,
17
+ "three_way_thresholds": {
18
+ "human_max": 0.5,
19
+ "ai_min": 0.8696600198745728,
20
+ "human_as_ai_validation_target": 0.05
21
+ }
22
+ }
inference.py ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import json
2
+ from pathlib import Path
3
+
4
+ import numpy as np
5
+ from peft import PeftModel
6
+ from safetensors.numpy import load_file
7
+ from sentence_transformers import SentenceTransformer
8
+
9
+
10
+ class MuncheAIDetector:
11
+ def __init__(self, model_dir: str | Path = "."):
12
+ model_dir = Path(model_dir)
13
+ self.encoder = SentenceTransformer("google/embeddinggemma-300m")
14
+ self.encoder[0].auto_model = PeftModel.from_pretrained(
15
+ self.encoder[0].auto_model,
16
+ model_dir,
17
+ )
18
+ self.encoder.max_seq_length = 2048
19
+ head = load_file(model_dir / "linear_head.safetensors")
20
+ self.weight = head["linear.weight"].reshape(-1)
21
+ self.bias = float(head["linear.bias"].item())
22
+ calibration = json.loads(
23
+ (model_dir / "calibration.json").read_text(encoding="utf-8")
24
+ )
25
+ self.human_max = calibration["thresholds"]["human_max"]
26
+ self.ai_min = calibration["thresholds"]["ai_min"]
27
+
28
+ def predict(self, text: str) -> dict:
29
+ tokens = self.encoder.tokenizer.encode(text, add_special_tokens=False)
30
+ if not 384 <= len(tokens) <= 2048:
31
+ raise ValueError("Input must contain 384 to 2,048 EmbeddingGemma tokens.")
32
+ embedding = self.encoder.encode(
33
+ [text],
34
+ normalize_embeddings=True,
35
+ convert_to_numpy=True,
36
+ )[0]
37
+ logit = float(embedding @ self.weight + self.bias)
38
+ score = (
39
+ float(1.0 / (1.0 + np.exp(-logit)))
40
+ if logit >= 0
41
+ else float(np.exp(logit) / (1.0 + np.exp(logit)))
42
+ )
43
+ label = (
44
+ "human"
45
+ if score <= self.human_max
46
+ else "llm"
47
+ if score >= self.ai_min
48
+ else "uncertain"
49
+ )
50
+ return {"label": label, "score": score, "tokens": len(tokens)}
length_stats.json ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "tokenizer": "google/embeddinggemma-300m",
3
+ "model_max_tokens": 2048,
4
+ "corpora": {
5
+ "independent_human": {
6
+ "samples": 783,
7
+ "raw_tokens": {
8
+ "minimum": 1708,
9
+ "p05": 1906,
10
+ "median": 2094,
11
+ "p95": 2258,
12
+ "maximum": 3920
13
+ },
14
+ "effective_tokens": {
15
+ "minimum": 1708,
16
+ "p05": 1906,
17
+ "median": 2048,
18
+ "p95": 2048,
19
+ "maximum": 2048
20
+ },
21
+ "truncated_fraction": 0.7075351213282248
22
+ },
23
+ "independent_ai": {
24
+ "samples": 330,
25
+ "raw_tokens": {
26
+ "minimum": 891,
27
+ "p05": 1271,
28
+ "median": 1999,
29
+ "p95": 2333,
30
+ "maximum": 2425
31
+ },
32
+ "effective_tokens": {
33
+ "minimum": 891,
34
+ "p05": 1271,
35
+ "median": 1999,
36
+ "p95": 2048,
37
+ "maximum": 2048
38
+ },
39
+ "truncated_fraction": 0.4636363636363636
40
+ },
41
+ "rewrite_human": {
42
+ "samples": 1306,
43
+ "raw_tokens": {
44
+ "minimum": 377,
45
+ "p05": 390,
46
+ "median": 406,
47
+ "p95": 430,
48
+ "maximum": 785
49
+ },
50
+ "effective_tokens": {
51
+ "minimum": 377,
52
+ "p05": 390,
53
+ "median": 406,
54
+ "p95": 430,
55
+ "maximum": 785
56
+ },
57
+ "truncated_fraction": 0.0
58
+ },
59
+ "rewrite_ai": {
60
+ "samples": 1306,
61
+ "raw_tokens": {
62
+ "minimum": 240,
63
+ "p05": 326,
64
+ "median": 420,
65
+ "p95": 514,
66
+ "maximum": 680
67
+ },
68
+ "effective_tokens": {
69
+ "minimum": 240,
70
+ "p05": 326,
71
+ "median": 420,
72
+ "p95": 514,
73
+ "maximum": 680
74
+ },
75
+ "truncated_fraction": 0.0
76
+ },
77
+ "sol_medium": {
78
+ "samples": 30,
79
+ "raw_tokens": {
80
+ "minimum": 1737,
81
+ "p05": 1974,
82
+ "median": 2173,
83
+ "p95": 2379,
84
+ "maximum": 2474
85
+ },
86
+ "effective_tokens": {
87
+ "minimum": 1737,
88
+ "p05": 1974,
89
+ "median": 2048,
90
+ "p95": 2048,
91
+ "maximum": 2048
92
+ },
93
+ "truncated_fraction": 0.8333333333333334
94
+ }
95
+ }
96
+ }
linear_head.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3c5e32168ee1bd4fabf22fbc3914c78e580bb420fec630b900690ab9f582e42e
3
+ size 3220
metrics.json ADDED
@@ -0,0 +1,146 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "Munche-768-AI-Detector",
3
+ "selected_step": 275,
4
+ "splits": {
5
+ "train": {
6
+ "human": 1315,
7
+ "llm": 1135
8
+ },
9
+ "validation": {
10
+ "human": 290,
11
+ "llm": 275
12
+ },
13
+ "test": {
14
+ "human": 277,
15
+ "llm": 256
16
+ }
17
+ },
18
+ "binary_test": {
19
+ "combined": {
20
+ "auroc": 0.9859403203971119,
21
+ "balanced_accuracy": 0.9391287793321299,
22
+ "human_recall": 0.9602888086642599,
23
+ "ai_recall": 0.91796875,
24
+ "confusion_matrix_human_ai": [
25
+ [
26
+ 266,
27
+ 11
28
+ ],
29
+ [
30
+ 21,
31
+ 235
32
+ ]
33
+ ]
34
+ },
35
+ "rewrite": {
36
+ "auroc": 0.9746814404432134,
37
+ "balanced_accuracy": 0.9157894736842105,
38
+ "human_recall": 0.9421052631578948,
39
+ "ai_recall": 0.8894736842105263,
40
+ "confusion_matrix_human_ai": [
41
+ [
42
+ 179,
43
+ 11
44
+ ],
45
+ [
46
+ 21,
47
+ 169
48
+ ]
49
+ ]
50
+ },
51
+ "independent": {
52
+ "auroc": 1.0,
53
+ "balanced_accuracy": 1.0,
54
+ "human_recall": 1.0,
55
+ "ai_recall": 1.0,
56
+ "confusion_matrix_human_ai": [
57
+ [
58
+ 87,
59
+ 0
60
+ ],
61
+ [
62
+ 0,
63
+ 66
64
+ ]
65
+ ]
66
+ }
67
+ },
68
+ "three_way": {
69
+ "selection": {
70
+ "split": "validation",
71
+ "human_as_ai_max": 0.05,
72
+ "human_boundary": "fixed at the binary threshold"
73
+ },
74
+ "thresholds": {
75
+ "human_max": 0.5,
76
+ "ai_min": 0.8696600198745728
77
+ },
78
+ "validation": {
79
+ "combined": {
80
+ "samples": 565,
81
+ "coverage": 0.9079646017699115,
82
+ "covered_accuracy": 0.9551656920077972,
83
+ "human_as_ai": 0.04827586206896552,
84
+ "human_uncertain": 0.08275862068965517,
85
+ "ai_as_human": 0.03272727272727273,
86
+ "ai_uncertain": 0.10181818181818182
87
+ },
88
+ "rewrite": {
89
+ "samples": 406,
90
+ "coverage": 0.8768472906403941,
91
+ "covered_accuracy": 0.9353932584269663,
92
+ "human_as_ai": 0.06896551724137931,
93
+ "human_uncertain": 0.11330049261083744,
94
+ "ai_as_human": 0.04433497536945813,
95
+ "ai_uncertain": 0.1330049261083744
96
+ },
97
+ "independent": {
98
+ "samples": 153,
99
+ "coverage": 0.9934640522875817,
100
+ "covered_accuracy": 1.0,
101
+ "human_as_ai": 0.0,
102
+ "human_uncertain": 0.011494252873563218,
103
+ "ai_as_human": 0.0,
104
+ "ai_uncertain": 0.0
105
+ },
106
+ "sol_medium": {
107
+ "samples": 6,
108
+ "coverage": 0.8333333333333334,
109
+ "covered_accuracy": 1.0,
110
+ "human_as_ai": null,
111
+ "human_uncertain": null,
112
+ "ai_as_human": 0.0,
113
+ "ai_uncertain": 0.16666666666666666
114
+ }
115
+ },
116
+ "test": {
117
+ "combined": {
118
+ "samples": 533,
119
+ "coverage": 0.9362101313320825,
120
+ "covered_accuracy": 0.9519038076152304,
121
+ "human_as_ai": 0.010830324909747292,
122
+ "human_uncertain": 0.02888086642599278,
123
+ "ai_as_human": 0.08203125,
124
+ "ai_uncertain": 0.1015625
125
+ },
126
+ "rewrite": {
127
+ "samples": 380,
128
+ "coverage": 0.9157894736842105,
129
+ "covered_accuracy": 0.9310344827586207,
130
+ "human_as_ai": 0.015789473684210527,
131
+ "human_uncertain": 0.042105263157894736,
132
+ "ai_as_human": 0.11052631578947368,
133
+ "ai_uncertain": 0.12631578947368421
134
+ },
135
+ "independent": {
136
+ "samples": 153,
137
+ "coverage": 0.9869281045751634,
138
+ "covered_accuracy": 1.0,
139
+ "human_as_ai": 0.0,
140
+ "human_uncertain": 0.0,
141
+ "ai_as_human": 0.0,
142
+ "ai_uncertain": 0.030303030303030304
143
+ }
144
+ }
145
+ }
146
+ }