Remidesbois commited on
Commit
7d7b358
·
verified ·
1 Parent(s): 4418169

Add files using upload-large-folder tool

Browse files
README.md CHANGED
@@ -1,27 +1,142 @@
1
  ---
2
- license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
4
 
5
  # Surya Bubble OCR Poneglyph
6
 
7
- Fine-tuned `datalab-to/surya-ocr-2` crop OCR model for Projet Poneglyph manga
8
- bubble transcription.
 
 
9
 
10
- ## Latest Modal H100 Benchmark
11
 
12
- Run completed on 2026-07-01 from training job
13
- `229565d7-0dd2-426f-8823-09e5812635a1`.
14
 
15
- | Metric | Value |
 
 
 
 
 
 
 
 
16
  | --- | ---: |
17
- | CER | 2.028% |
18
- | WER | 3.078% |
19
- | Exact match | 90.09% |
20
- | Average Levenshtein | 0.667 chars |
21
- | Blank rate | 0.00% |
22
- | Multiline rate | 0.00% |
23
- | Hallucination rate | 0.00% |
24
- | Test samples | 1221 |
25
-
26
- The model is active for `surya_bubble_ocr` in the Poneglyph training registry,
27
- even though the latest LightOn crop OCR benchmark is stronger.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: openrail
3
+ base_model: datalab-to/surya-ocr-2
4
+ library_name: transformers
5
+ pipeline_tag: image-text-to-text
6
+ language:
7
+ - fr
8
+ tags:
9
+ - ocr
10
+ - manga
11
+ - comics
12
+ - qwen3.5
13
+ - surya
14
+ model-index:
15
+ - name: Surya Bubble OCR Poneglyph
16
+ results:
17
+ - task:
18
+ type: image-text-to-text
19
+ name: Bubble text transcription
20
+ dataset:
21
+ name: Poneglyph held-out bubble OCR
22
+ type: custom
23
+ split: test
24
+ metrics:
25
+ - type: cer
26
+ value: 0.0045099636
27
+ name: CER
28
+ - type: wer
29
+ value: 0.0165559530
30
+ name: WER
31
+ - type: exact_match
32
+ value: 0.9065354884
33
+ name: Exact match
34
  ---
35
 
36
  # Surya Bubble OCR Poneglyph
37
 
38
+ Fine-tune de [`datalab-to/surya-ocr-2`](https://huggingface.co/datalab-to/surya-ocr-2)
39
+ pour la transcription exacte de bulles de manga francophones recadrées.
40
+ Ce modèle transcrit une bulle à la fois ; il ne détecte pas les bulles et ne
41
+ renvoie pas de bounding boxes.
42
 
43
+ ## Résultats
44
 
45
+ Entraînement local sur une NVIDIA RTX 3090. Le split est effectué par page :
46
+ aucune page source n'est partagée entre train, validation et test.
47
 
48
+ | Split | Pages | Bulles |
49
+ | --- | ---: | ---: |
50
+ | Train | 749 | 6 793 |
51
+ | Validation | 161 | 1 311 |
52
+ | Test held-out | 161 | 1 423 |
53
+
54
+ Benchmark final exhaustif sur les 1 423 bulles du test held-out :
55
+
56
+ | Métrique | Résultat |
57
  | --- | ---: |
58
+ | CER | **0,451 %** |
59
+ | WER | **1,656 %** |
60
+ | Exact match | **90,65 %** |
61
+ | Levenshtein moyen | **0,1595 caractère** |
62
+ | Sorties vides | **0 / 1 423** |
63
+ | Hallucinations sur références vides | **0** |
64
+ | Limite de génération atteinte | **0 / 1 423** |
65
+
66
+ Les textes très courts de 1 à 2 caractères atteignent 100 % d'exact match
67
+ sur 98 exemples. Les erreurs restantes se concentrent principalement sur les
68
+ onomatopées ambiguës, la casse et les répétitions de rires.
69
+
70
+ Le fichier [`benchmark_test.json`](benchmark_test.json) contient les métriques,
71
+ les tranches par longueur et les prédictions de chaque échantillon.
72
+
73
+ ## Utilisation
74
+
75
+ ```python
76
+ import torch
77
+ from PIL import Image
78
+ from transformers import AutoModelForImageTextToText, AutoProcessor
79
+
80
+ model_id = "Remidesbois/surya-bubble-ocr-poneglyph"
81
+ processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
82
+ model = AutoModelForImageTextToText.from_pretrained(
83
+ model_id,
84
+ dtype=torch.bfloat16,
85
+ device_map="cuda",
86
+ trust_remote_code=True,
87
+ ).eval()
88
+
89
+ image = Image.open("bulle.png").convert("RGB")
90
+ messages = [{
91
+ "role": "user",
92
+ "content": [
93
+ {"type": "image", "image": "bulle.png"},
94
+ {
95
+ "type": "text",
96
+ "text": "Transcris exactement le texte visible dans cette bulle. Ne rajoute rien.",
97
+ },
98
+ ],
99
+ }]
100
+ prompt = processor.apply_chat_template(
101
+ messages,
102
+ add_generation_prompt=True,
103
+ tokenize=False,
104
+ )
105
+ inputs = processor(text=[prompt], images=[image], return_tensors="pt").to("cuda")
106
+
107
+ with torch.inference_mode():
108
+ output_ids = model.generate(
109
+ **inputs,
110
+ max_new_tokens=256,
111
+ do_sample=False,
112
+ )
113
+
114
+ prompt_tokens = inputs["input_ids"].shape[1]
115
+ text = processor.batch_decode(
116
+ output_ids[:, prompt_tokens:],
117
+ skip_special_tokens=True,
118
+ )[0].strip()
119
+ print(text)
120
+ ```
121
+
122
+ ## Entraînement
123
+
124
+ - 665,7 M paramètres, dont 606,0 M entraînables ;
125
+ - modèle langage complet, merger multimodal et 4 derniers blocs vision ;
126
+ - BF16 et TF32 ;
127
+ - batch physique 16, accumulation de gradient 2 ;
128
+ - 5 époques, sélection du meilleur checkpoint sur le CER génératif ;
129
+ - budget de génération de 256 tokens.
130
+
131
+ Le pipeline reproductible se trouve dans le dossier
132
+ `docker_scripts/finetune_surya_bubble_ocr` du projet Poneglyph.
133
+
134
+ ## Limites
135
+
136
+ - Le test est un holdout par page issu du même projet et du même processus de
137
+ validation que le train ; il ne mesure pas une généralisation universelle à
138
+ tous les mangas, langues ou styles d'impression.
139
+ - Le modèle attend un crop contenant une seule zone de texte.
140
+ - Les onomatopées rares ou très stylisées restent la principale source
141
+ d'erreurs.
142
+ - La licence `openrail` est héritée du modèle de base.
benchmark_test.json CHANGED
The diff for this file is too large to render. See raw diff
 
config.json CHANGED
@@ -78,7 +78,7 @@
78
  "vocab_size": 65425
79
  },
80
  "tie_word_embeddings": true,
81
- "transformers_version": "5.13.0.dev0",
82
  "use_cache": true,
83
  "video_token_id": 12,
84
  "vision_config": {
 
78
  "vocab_size": 65425
79
  },
80
  "tie_word_embeddings": true,
81
+ "transformers_version": "5.14.1",
82
  "use_cache": true,
83
  "video_token_id": 12,
84
  "vision_config": {
generation_config.json CHANGED
@@ -2,8 +2,8 @@
2
  "_from_model_config": true,
3
  "do_sample": false,
4
  "eos_token_id": 2,
5
- "max_new_tokens": 96,
6
  "pad_token_id": 0,
7
- "transformers_version": "5.13.0.dev0",
8
  "use_cache": true
9
  }
 
2
  "_from_model_config": true,
3
  "do_sample": false,
4
  "eos_token_id": 2,
5
+ "max_new_tokens": 256,
6
  "pad_token_id": 0,
7
+ "transformers_version": "5.14.1",
8
  "use_cache": true
9
  }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:7e942d090963621481f202af702282172b9fac5534df2d9530eda172807fb777
3
  size 1331461328
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d3554f1487caa1fd33c30e6a397d306b0d71fe041639f8668482ad3b02a223d
3
  size 1331461328
processor_config.json CHANGED
@@ -22,7 +22,7 @@
22
  "resample": 3,
23
  "rescale_factor": 0.00392156862745098,
24
  "size": {
25
- "longest_edge": 16777216,
26
  "shortest_edge": 65536
27
  },
28
  "temporal_patch_size": 2
 
22
  "resample": 3,
23
  "rescale_factor": 0.00392156862745098,
24
  "size": {
25
+ "longest_edge": 1048576,
26
  "shortest_edge": 65536
27
  },
28
  "temporal_patch_size": 2
tokenizer.json CHANGED
@@ -1,7 +1,14 @@
1
  {
2
  "version": "1.0",
3
  "truncation": null,
4
- "padding": null,
 
 
 
 
 
 
 
5
  "added_tokens": [
6
  {
7
  "id": 0,
 
1
  {
2
  "version": "1.0",
3
  "truncation": null,
4
+ "padding": {
5
+ "strategy": "BatchLongest",
6
+ "direction": "Left",
7
+ "pad_to_multiple_of": 16,
8
+ "pad_id": 0,
9
+ "pad_type_id": 0,
10
+ "pad_token": "<|endoftext|>"
11
+ },
12
  "added_tokens": [
13
  {
14
  "id": 0,