fastx-ai
/

Marco-o1-1.2B-mlx-int4

Text Generation

text-generation-inference

4-bit precision

Model card Files Files and versions Community

starkx commited on Nov 30, 2024

Commit

92e0c3c

·

verified ·

1 Parent(s): d219b08

Update README.md

Files changed (1) hide show

README.md +28 -0

README.md CHANGED Viewed

@@ -63,3 +63,31 @@ if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not
 response = generate(model, tokenizer, prompt=prompt, verbose=True)
 ```

 response = generate(model, tokenizer, prompt=prompt, verbose=True)
 ```
+## Change system prompt ...
+1. clone this repo to local
+2. change tokenizer_config.json
+```json
+"chat_template": "{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system\n\n你是一个经过良好训练的AI助手，你的名字是Marco-o1.\n        \n## 重要！！！！！\n当你回答问题时，你的思考应该在<Thought>内完成，<Output>内输出你的结果。\n<Thought>应该尽可能是英文，但是有2个特例，一个是对原文中的引用，另一个是是数学应该使用markdown格式，<Output>内的输出需要遵循用户输入的语言。\n        <|im_end|>\n' }}{% endif %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}",
+```
+3. load
+```python
+from mlx_lm import load, generate
+model, tokenizer = load("./mlx_model") # notice: folder where you put this repo files.
+prompt="hello, can you teach me why 2 + 4 = 6 ?"
+if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
+    messages = [{"role": "user", "content": prompt}]
+    prompt = tokenizer.apply_chat_template(
+        messages, tokenize=False, add_generation_prompt=True
+    )
+response = generate(model, tokenizer, prompt=prompt, verbose=True)
+```