File size: 9,490 Bytes
9bffc08 76f1c0b 9bffc08 c3a152c e38399b b68784f 9bffc08 3ddf676 b59762c 1c01c7d 90d8304 f15e06a 080e5cf 89970fd 859c225 575f672 8d389d9 b041454 aa5b29f 4789442 921f867 3512733 2d3aaf9 c9406da 9e8aaf6 a91496b 9bffc08 82a69a2 9bffc08 |
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 |
---
base_model: microsoft/Phi-3-medium-4k-instruct
inference: false
language:
- multilingual
library_name: gguf
license: mit
license_link: https://huggingface.co./microsoft/Phi-3-medium-4k-instruct/resolve/main/LICENSE
pipeline_tag: text-generation
quantized_by: legraphista
tags:
- quantized
- GGUF
- imatrix
- quantization
- imat
- imatrix
- static
---
# Phi-3-medium-4k-instruct-IMat-GGUF
_Llama.cpp imatrix quantization of microsoft/Phi-3-medium-4k-instruct_
Original Model: [microsoft/Phi-3-medium-4k-instruct](https://huggingface.co./microsoft/Phi-3-medium-4k-instruct)
Original dtype: `BF16` (`bfloat16`)
Quantized by: llama.cpp [b2998](https://github.com/ggerganov/llama.cpp/releases/tag/b2998)
IMatrix dataset: [here](https://gist.githubusercontent.com/legraphista/d6d93f1a254bcfc58e0af3777eaec41e/raw/d380e7002cea4a51c33fffd47db851942754e7cc/imatrix.calibration.medium.raw)
## Files
### IMatrix
Status: β
Available
Link: [here](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/imatrix.dat)
### Common Quants
| Filename | Quant type | File Size | Status | Uses IMatrix | Is Split |
| -------- | ---------- | --------- | ------ | ------------ | -------- |
| [Phi-3-medium-4k-instruct.Q8_0.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q8_0.gguf) | Q8_0 | 14.83GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.Q6_K.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q6_K.gguf) | Q6_K | 11.45GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.Q4_K.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q4_K.gguf) | Q4_K | 8.57GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.Q3_K.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q3_K.gguf) | Q3_K | 6.92GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.Q2_K.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q2_K.gguf) | Q2_K | 5.14GB | β
Available | π’ Yes | π¦ No
### All Quants
| Filename | Quant type | File Size | Status | Uses IMatrix | Is Split |
| -------- | ---------- | --------- | ------ | ------------ | -------- |
| [Phi-3-medium-4k-instruct.FP16.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.FP16.gguf) | F16 | 27.92GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.BF16.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.BF16.gguf) | BF16 | 27.92GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.Q5_K.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q5_K.gguf) | Q5_K | 10.07GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.Q5_K_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q5_K_S.gguf) | Q5_K_S | 9.62GB | β
Available | βͺ No | π¦ No
| [Phi-3-medium-4k-instruct.Q4_K_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q4_K_S.gguf) | Q4_K_S | 7.95GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.Q3_K_L.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q3_K_L.gguf) | Q3_K_L | 7.49GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.Q3_K_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q3_K_S.gguf) | Q3_K_S | 6.06GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.Q2_K_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.Q2_K_S.gguf) | Q2_K_S | 4.77GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ4_NL.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ4_NL.gguf) | IQ4_NL | 7.90GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ4_XS.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ4_XS.gguf) | IQ4_XS | 7.47GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ3_M.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ3_M.gguf) | IQ3_M | 6.47GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ3_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ3_S.gguf) | IQ3_S | 6.06GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ3_XS.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ3_XS.gguf) | IQ3_XS | 5.81GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ3_XXS.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ3_XXS.gguf) | IQ3_XXS | 5.45GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ2_M.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ2_M.gguf) | IQ2_M | 4.72GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ2_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ2_S.gguf) | IQ2_S | 4.34GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ2_XS.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ2_XS.gguf) | IQ2_XS | 4.13GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ2_XXS.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ2_XXS.gguf) | IQ2_XXS | 3.72GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ1_M.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ1_M.gguf) | IQ1_M | 3.24GB | β
Available | π’ Yes | π¦ No
| [Phi-3-medium-4k-instruct.IQ1_S.gguf](https://huggingface.co./legraphista/Phi-3-medium-4k-instruct-IMat-GGUF/blob/main/Phi-3-medium-4k-instruct.IQ1_S.gguf) | IQ1_S | 2.96GB | β
Available | π’ Yes | π¦ No
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download legraphista/Phi-3-medium-4k-instruct-IMat-GGUF --include "Phi-3-medium-4k-instruct.Q8_0.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download legraphista/Phi-3-medium-4k-instruct-IMat-GGUF --include "Phi-3-medium-4k-instruct.Q8_0/*" --local-dir Phi-3-medium-4k-instruct.Q8_0
# see FAQ for merging GGUF's
```
---
## Inference
### Simple chat template
```
<s><|user|>
I am going to Paris, what should I see?<|end|>
<|assistant|>
Paris, the capital of France, is known for its stunning architecture, art museums, historical landmarks, and romantic atmosphere. Here are some of the top attractions to see in Paris:
1. The Eiffel Tower: The iconic Eiffel Tower is one of the most recognizable landmarks in the world and offers breathtaking views of the city.
2. The Louvre Museum: The Louvre is one of the world's largest and most famous museums, housing an impressive collection of art and artifacts, including the Mona Lisa.
3. Notre-Dame Cathedral: This beautiful cathedral is one of the most famous landmarks in Paris and is known for its Gothic architecture and stunning stained glass windows.
These are just a few of the many attractions that Paris has to offer. With so much to see and do, it's no wonder that Paris is one of the most popular tourist destinations in the world."<|end|>
<|user|>
What is so great about #1?<|end|>
<|assistant|>
```
### Llama.cpp
```
llama.cpp/main -m Phi-3-medium-4k-instruct.Q8_0.gguf --color -i -p "prompt here (according to the chat template)"
```
---
## FAQ
### Why is the IMatrix not applied everywhere?
According to [this investigation](https://www.reddit.com/r/LocalLLaMA/comments/1993iro/ggufs_quants_can_punch_above_their_weights_now/), it appears that lower quantizations are the only ones that benefit from the imatrix input (as per hellaswag results).
### How do I merge a split GGUF?
1. Make sure you have `gguf-split` available
- To get hold of `gguf-split`, navigate to https://github.com/ggerganov/llama.cpp/releases
- Download the appropriate zip for your system from the latest release
- Unzip the archive and you should be able to find `gguf-split`
2. Locate your GGUF chunks folder (ex: `Phi-3-medium-4k-instruct.Q8_0`)
3. Run `gguf-split --merge Phi-3-medium-4k-instruct.Q8_0/Phi-3-medium-4k-instruct.Q8_0-00001-of-XXXXX.gguf Phi-3-medium-4k-instruct.Q8_0.gguf`
- Make sure to point `gguf-split` to the first chunk of the split.
---
Got a suggestion? Ping me [@legraphista](https://x.com/legraphista)! |