Hugging Face
Models
Datasets
Spaces
Posts
Docs
Solutions
Pricing
Log In
Sign Up
line-corporation
/
sacpo
like
5
Follow
LINE
26
Reinforcement Learning
Transformers
Safetensors
PKU-Alignment/PKU-SafeRLHF-30K
English
llama
text-generation
reinforcement-learning-from-human-feedback
rlhf
safety
ai-safety
alpaca
text-generation-inference
Inference Endpoints
arxiv:
2404.11049
arxiv:
2305.18290
License:
cc-by-nc-4.0
Model card
Files
Files and versions
Community
Train
Deploy
Use this model
main
sacpo
Commit History
Update README.md
b596248
verified
akifumiwachi
commited on
Jun 21
Update README.md
15ddeb1
verified
reisato80
commited on
Jun 19
Upload LlamaForCausalLM
06ea01f
verified
reisato80
commited on
Jun 19
Upload tokenizer
e8f6e19
verified
reisato80
commited on
Jun 19
Create README.md
c4c56bb
verified
reisato80
commited on
Jun 19
initial commit
221953a
verified
ospo-line
commited on
Jun 11