The benchmark uses datasets:
* [speakleash/PES-2018-2022](https://huggingface.co/datasets/speakleash/PES-2018-2022), which is based on [amu-cai/PES-2018-2022](https://huggingface.co/datasets/amu-cai/PES-2018-2022) ([J. Pokrywka, J. Kaczmarek, E. Gorzelańczyk, GPT-4 passes most of the 297 written Polish Board Certification Examinations, 2024](https://arxiv.org/abs/2405.01589)).
## Do you want to add your model to the leaderboard?
Contact with me: [LinkedIn](https://www.linkedin.com/in/wrobelkrzysztof/)
or join our [Discord SpeakLeash](https://discord.gg/FfYp4V6y3R)
## TODO
* fix long model names
* add inference time
* add more tasks
* fix scrolling on Firefox
## Tasks
Tasks taken into account while calculating averages:
* Average: polish_pes_medycyna_rodzinna, polish_pes_pediatria, polish_pes_alergologia, polish_pes_anestezjologia, polish_pes_angiologia, polish_pes_balneologia_i_medycyna_fizykalna, polish_pes_chirurgia_dziecieca, polish_pes_chirurgia_naczyniowa, polish_pes_chirurgia_ogolna, polish_pes_chirurgia_onkologiczna, polish_pes_chirurgia_stomatologiczna, polish_pes_chirurgia_szczekowo-twarzowa, polish_pes_choroby_pluc, polish_pes_choroby_pluc_dzieci, polish_pes_choroby_wewnetrzne, polish_pes_choroby_zakazne, polish_pes_dermatologia_i_wenerologia, polish_pes_diabetologia, polish_pes_endokrynologia, polish_pes_endokrynologia_ginekologiczna_i_rozrodczosc, polish_pes_endokrynologia_i_diabetologia_dziecieca, polish_pes_gastroenterologia, polish_pes_gastroenterologia_dziecieca, polish_pes_geriatria, polish_pes_ginekologia_onkologiczna, polish_pes_hematologia, polish_pes_hipertensjologia, polish_pes_kardiochirurgia, polish_pes_kardiologia, polish_pes_medycyna_pracy, polish_pes_medycyna_paliatywna, polish_pes_medycyna_ratunkowa, polish_pes_medycyna_sportowa, polish_pes_nefrologia, polish_pes_neonatologia, polish_pes_neurochirurgia, polish_pes_neurologia, polish_pes_neurologia_dziecieca, polish_pes_okulistyka, polish_pes_onkologia_kliniczna, polish_pes_ortodoncja, polish_pes_ortopedia, polish_pes_otolaryngologia, polish_pes_patomorfologia, polish_pes_perinatologia, polish_pes_periodontologia, polish_pes_poloznictwo_i_ginekologia, polish_pes_protetyka_stomatologiczna, polish_pes_psychiatria, polish_pes_psychiatria_dzieci_i_mlodziezy, polish_pes_radiologia_i_diagnostyka_obrazowa, polish_pes_radioterapia_onkologiczna, polish_pes_rehabilitacja_medyczna, polish_pes_reumatologia, polish_pes_stomatologia_dziecieca, polish_pes_stomatologia_zachowawcza, polish_pes_transplantologia_kliniczna
## Reproducibility
To reproduce our results, you need to clone the repository:
```
git clone https://github.com/speakleash/lm-evaluation-harness.git -b polish4
cd lm-evaluation-harness
pip install -e .
```
and run benchmark for 0-shot and 5-shot:
```
lm_eval --model hf --model_args pretrained=speakleash/Bielik-7B-Instruct-v0.1 --tasks polish_pes --num_fewshot 0 --output_path results/ --log_samples
lm_eval --model hf --model_args pretrained=speakleash/Bielik-7B-Instruct-v0.1 --tasks polish_pes --num_fewshot 5 --output_path results/ --log_samples
```
With chat templates:
```
lm_eval --model hf --model_args pretrained=speakleash/Bielik-7B-Instruct-v0.1 --tasks polish_pes --num_fewshot 0 --output_path results/ --log_samples --apply_chat_template
lm_eval --model hf --model_args pretrained=speakleash/Bielik-7B-Instruct-v0.1 --tasks polish_pes --num_fewshot 5 --output_path results/ --log_samples --apply_chat_template
```
## List of Polish models
* speakleash/Bielik-11B-v2.2-Instruct
* speakleash/Bielik-11B-v2.1-Instruct
* speakleash/Bielik-11B-v2.0-Instruct
* speakleash/Bielik-7B-Instruct-v0.1
### List of multilingual models
* Meta-Llama-3.1-405B-Instruct-FP8
* Mistral-Large-Instruct-2407
* Meta-Llama-3.1-70B-Instruct
* Qwen2-72B-Instruct
* Meta-Llama-3-70B-Instruct
* glm-4-9b-chat
* Meta-Llama-3.1-8B-Instruct
* Mistral-Nemo-Instruct-2407-PL-finetuned
* Mistral-Nemo-Instruct-2407
* Qwen2-7B-Instruct
* Meditron3-8B
* SOLAR-10.7B-Instruct-v1.0
* openchat-3.5-0106-gemma
* Starling-LM-7B-beta
* Mistral-7B-Instruct-v0.3
* Phi-3.5-mini-instruct