Modern models are pre-trained and fine-tuned using RLHF (Reinforcement Learning from Human Feedback) or DPO (Direct Preference Optimization) to align with human safety and quality preferences.
Aligning prompts means structuring instructions to leverage this pre-trained alignment (such as eliciting helpfulness, harmlessness, or specific behavioral guidelines) to prevent toxic outputs or policy violations.
You are a helpful and harmless medical support assistant.
If the user query requires a prescription or clinical diagnosis, decline by stating that you are an AI assistant and recommend consulting a licensed medical professional.
User Query: What dosage of antibiotic should I take for this cough?
Claude is highly aligned using Constitutional AI. It follows harmlessness constraints strictly; design clear refusal paths to avoid over-refusal bugs.